> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fluffbuzz.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Inferrs

[inferrs](https://github.com/ericcurtin/inferrs) can serve local models behind an
OpenAI-compatible `/v1` API. FluffBuzz works with `inferrs` through the generic
`openai-completions` path.

`inferrs` is currently best treated as a custom self-hosted OpenAI-compatible
backend, not a dedicated FluffBuzz provider plugin.

## Getting started

<Steps>
  <Step title="Start inferrs with a model">
    ```bash theme={"theme":{"light":"min-light","dark":"min-dark"}}
    inferrs serve google/gemma-4-E2B-it \
      --host 127.0.0.1 \
      --port 8080 \
      --device metal
    ```
  </Step>

  <Step title="Verify the server is reachable">
    ```bash theme={"theme":{"light":"min-light","dark":"min-dark"}}
    curl http://127.0.0.1:8080/health
    curl http://127.0.0.1:8080/v1/models
    ```
  </Step>

  <Step title="Add an FluffBuzz provider entry">
    Add an explicit provider entry and point your default model at it. See the full config example below.
  </Step>
</Steps>

## Full config example

This example uses Gemma 4 on a local `inferrs` server.

```json5 theme={"theme":{"light":"min-light","dark":"min-dark"}}
{
  agents: {
    defaults: {
      model: { primary: "inferrs/google/gemma-4-E2B-it" },
      models: {
        "inferrs/google/gemma-4-E2B-it": {
          alias: "Gemma 4 (inferrs)",
        },
      },
    },
  },
  models: {
    mode: "merge",
    providers: {
      inferrs: {
        baseUrl: "http://127.0.0.1:8080/v1",
        apiKey: "inferrs-local",
        api: "openai-completions",
        models: [
          {
            id: "google/gemma-4-E2B-it",
            name: "Gemma 4 E2B (inferrs)",
            reasoning: false,
            input: ["text"],
            cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
            contextWindow: 131072,
            maxTokens: 4096,
            compat: {
              requiresStringContent: true,
            },
          },
        ],
      },
    },
  },
}
```

## Advanced configuration

<AccordionGroup>
  <Accordion title="Why requiresStringContent matters">
    Some `inferrs` Chat Completions routes accept only string
    `messages[].content`, not structured content-part arrays.

    <Warning>
      If FluffBuzz runs fail with an error like:

      ```text theme={"theme":{"light":"min-light","dark":"min-dark"}}
      messages[1].content: invalid type: sequence, expected a string
      ```

      set `compat.requiresStringContent: true` in your model entry.
    </Warning>

    ```json5 theme={"theme":{"light":"min-light","dark":"min-dark"}}
    compat: {
      requiresStringContent: true
    }
    ```

    FluffBuzz will flatten pure text content parts into plain strings before sending
    the request.
  </Accordion>

  <Accordion title="Gemma and tool-schema caveat">
    Some current `inferrs` + Gemma combinations accept small direct
    `/v1/chat/completions` requests but still fail on full FluffBuzz agent-runtime
    turns.

    If that happens, try this first:

    ```json5 theme={"theme":{"light":"min-light","dark":"min-dark"}}
    compat: {
      requiresStringContent: true,
      supportsTools: false
    }
    ```

    That disables FluffBuzz's tool schema surface for the model and can reduce prompt
    pressure on stricter local backends.

    If tiny direct requests still work but normal FluffBuzz agent turns continue to
    crash inside `inferrs`, the remaining issue is usually upstream model/server
    behavior rather than FluffBuzz's transport layer.
  </Accordion>

  <Accordion title="Manual smoke test">
    Once configured, test both layers:

    ```bash theme={"theme":{"light":"min-light","dark":"min-dark"}}
    curl http://127.0.0.1:8080/v1/chat/completions \
      -H 'content-type: application/json' \
      -d '{"model":"google/gemma-4-E2B-it","messages":[{"role":"user","content":"What is 2 + 2?"}],"stream":false}'
    ```

    ```bash theme={"theme":{"light":"min-light","dark":"min-dark"}}
    fluffbuzz infer model run \
      --model inferrs/google/gemma-4-E2B-it \
      --prompt "What is 2 + 2? Reply with one short sentence." \
      --json
    ```

    If the first command works but the second fails, check the troubleshooting section below.
  </Accordion>

  <Accordion title="Proxy-style behavior">
    `inferrs` is treated as a proxy-style OpenAI-compatible `/v1` backend, not a
    native OpenAI endpoint.

    * Native OpenAI-only request shaping does not apply here
    * No `service_tier`, no Responses `store`, no prompt-cache hints, and no
      OpenAI reasoning-compat payload shaping
    * Hidden FluffBuzz attribution headers (`originator`, `version`, `User-Agent`)
      are not injected on custom `inferrs` base URLs
  </Accordion>
</AccordionGroup>

## Troubleshooting

<AccordionGroup>
  <Accordion title="curl /v1/models fails">
    `inferrs` is not running, not reachable, or not bound to the expected
    host/port. Make sure the server is started and listening on the address you
    configured.
  </Accordion>

  <Accordion title="messages[].content expected a string">
    Set `compat.requiresStringContent: true` in the model entry. See the
    `requiresStringContent` section above for details.
  </Accordion>

  <Accordion title="Direct /v1/chat/completions calls pass but fluffbuzz infer model run fails">
    Try setting `compat.supportsTools: false` to disable the tool schema surface.
    See the Gemma tool-schema caveat above.
  </Accordion>

  <Accordion title="inferrs still crashes on larger agent turns">
    If FluffBuzz no longer gets schema errors but `inferrs` still crashes on larger
    agent turns, treat it as an upstream `inferrs` or model limitation. Reduce
    prompt pressure or switch to a different local backend or model.
  </Accordion>
</AccordionGroup>

<Tip>
  For general help, see [Troubleshooting](/help/troubleshooting) and [FAQ](/help/faq).
</Tip>

## Related

<CardGroup cols={2}>
  <Card title="Local models" href="/gateway/local-models" icon="server">
    Running FluffBuzz against local model servers.
  </Card>

  <Card title="Gateway troubleshooting" href="/gateway/troubleshooting#local-openai-compatible-backend-passes-direct-probes-but-agent-runs-fail" icon="wrench">
    Debugging local OpenAI-compatible backends that pass probes but fail agent runs.
  </Card>

  <Card title="Model selection" href="/concepts/model-providers" icon="layers">
    Overview of all providers, model refs, and failover behavior.
  </Card>
</CardGroup>
