Skip to content

Inference

BNW uses a small provider-neutral JSON format at the encrypted invocation boundary. Provider-local adapters translate it to OpenAI-compatible Chat Completions, Anthropic Messages, or future model APIs. Provider endpoints, credentials, model selection, timeouts, and routing policy are never requester-controlled fields.

{
"profile": "bnw.inference-request/1",
"system": "Answer accurately and concisely.",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What is the capital of France?"}
]
}
],
"generation": {
"max_output_tokens": 128,
"temperature": 0.2,
"top_p": 0.9,
"stop_sequences": ["END"]
}
}

profile and one or more messages are required. system and generation are optional. Version 1 supports user and assistant messages containing text blocks. temperature is between 0 and 1, top_p is greater than 0 and at most 1, and at most 16 stop sequences may be supplied. A requested output-token limit cannot exceed the model descriptor’s limit; when omitted, the descriptor limit is used. The Anthropic adapter requires one of those limits because its native API requires max_tokens.

A message may also carry images, for models whose descriptor lists image among its input modalities:

{"type": "image", "media_type": "image/png", "data": "<base64>"}
  • Limits. Images are PNG, JPEG, GIF, or WebP, and up to 5 MiB each once decoded.
  • Refusals. A provider refuses a request with images for a text-only model before accepting it, with a signed rejection.
  • Adapters.
    • The OpenAI-compatible adapter sends content parts with a data: URL, but only when a message contains an image. Text-only messages are sent as plain strings, as before.
    • The Anthropic adapter sends native image blocks.
  • Compatibility. Image blocks are an optional addition to version 1; older nodes reject requests that contain them.

A request may also offer the model tools, for providers whose offers list the inference-tools feature:

{
"profile": "bnw.inference-request/1",
"messages": [
{"role": "user", "content": [{"type": "text", "text": "Price of BNW?"}]},
{"role": "assistant", "content": [
{"type": "tool_call", "id": "call-1", "name": "t1_quotes", "input": {"ticker": "BNW"}}
]},
{"role": "user", "content": [
{"type": "tool_result", "call_id": "call-1", "content": "42"}
]}
],
"tools": [
{"name": "t1_quotes", "description": "Stock quotes", "input_schema": {"type": "object"}}
]
}
  • Tools. At most 16. Each has a unique name of 1-64 letters, digits, underscores, or dashes, an optional description, and a JSON Schema object of its input up to 16 KiB.
  • Calls and results. Only assistant messages carry tool_call blocks, and each names an offered tool with a unique ID and an input object up to 16 KiB. Only user messages carry tool_result blocks, each answering an earlier call, with up to 64 KiB of text and an optional is_error.
  • Responses. A response may contain tool_call blocks, with stop reason tool_use. The model only asks; BNW runs nothing. The requester decides what to call.
  • Adapters. The OpenAI-compatible adapter sends tools as functions, assistant calls as tool_calls, and each result as a tool message; it reads tool_calls back. The Anthropic adapter uses native tool_use and tool_result blocks.
  • Which backends take tools. A registration can say so (bnw inference register --tool-calling true|false). Unset, the Anthropic API and OpenAI-compatible endpoints with remote egress are assumed to, and local runtimes are not, since many small local models cannot. Sharing a model in one step follows what Ollama reports. A provider refuses a request with tools for a backend that does not take them, with a signed rejection, before anything runs.
  • Compatibility. Tools are an optional addition to version 1, like images. Older nodes reject requests that contain them, which is why requesters send them only to offers listing inference-tools.

The strict versioned format deliberately excludes provider model names, endpoints, credentials, streaming controls, retries, fallbacks, thinking controls, caching, and provider-specific extension objects. These can be added through later versioned contracts rather than silently changing version 1 semantics.

{
"profile": "bnw.inference-response/1",
"content": [
{"type": "text", "text": "Paris."}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 18,
"output_tokens": 3
},
"provider": {
"adapter": "anthropic-messages",
"runtime_model": "provider/runtime-model-id",
"request_id": "msg_example"
}
}

Normalized stop reasons are end_turn, max_output_tokens, stop_sequence, tool_use, refusal, and other. Usage, stop sequence, provider request ID, and even the provider-reported runtime model may be absent or inaccurate. They are non-authoritative metadata signed by the BNW provider, not proof of model execution or correctness.

Version 1 normalizes returned text and tool calls, and does not expose provider-specific reasoning, citation, or cache blocks. The reference adapter is non-streaming. Richer multimodal or structured-output contracts belong in later versions.

openai-chat translates the system prompt and messages to Chat Completions, inserts the fixed registered runtime model, forces stream: false, and maps the first choice into the BNW response. It is suitable for direct OpenAI-compatible services, Ollama, llama.cpp, vLLM, and an optional LiteLLM Proxy.

anthropic-messages translates the same request to the native Messages API, sends the API key using x-api-key, supplies the fixed anthropic-version, inserts the fixed registered runtime model, and normalizes the returned content, stop reason, usage, and IDs.

The public model_id in the capability descriptor and the provider’s runtime_model serve different purposes. The descriptor is the provider’s stable, discoverable claim. The local registration maps that claim to an Ollama name, hosted model identifier, deployment ID, or gateway alias. Both are provider-controlled and fixed before an invocation executes. The signed calling progress records the mapping, while the normalized result records the model string returned by the runtime; none is cryptographic proof that particular weights ran.

The local registry records payload_egress separately from the URL. local is accepted only for a loopback endpoint. remote is required whenever plaintext ultimately leaves the host, including when BNW calls a loopback LiteLLM proxy that forwards to a hosted provider. BNW cannot inspect a gateway’s downstream routing, logging, caching, retry, or fallback configuration; the provider remains responsible for configuring those behaviors consistently with its published capability.

Hosted routers: OpenRouter and Hugging Face

Section titled “Hosted routers: OpenRouter and Hugging Face”

--preset openrouter and --preset huggingface set up the OpenAI-compatible adapter for a hosted router in one step. Explicit options still win.

Preset Endpoint Key secret Model names
openrouter https://openrouter.ai/api/v1/chat/completions OPENROUTER_API_KEY vendor/model, as OpenRouter lists them
huggingface https://router.huggingface.co/v1/chat/completions HF_TOKEN org/model, or org/model:provider to pin one inference provider
Terminal window
bnw provider secret set OPENROUTER_API_KEY
bnw inference models openrouter --search claude --tools
bnw inference register <CAPABILITY_ID> --preset openrouter --runtime-model anthropic/claude-opus-5.5
  • Payload egress. A preset sends prompts to the service, so egress is remote.
  • Headers. OpenRouter also gets its app-attribution headers (HTTP-Referer, X-Title). --header NAME=VALUE adds any other non-secret header, such as X-HF-Bill-To. Extra headers cannot replace the adapter’s own Authorization, x-api-key, anthropic-version, or Accept, and they cannot carry control characters.
  • Model lists. bnw inference models reads the router’s public model list: context size, output limit, whether the model takes tools, and price per million tokens (for Hugging Face, the cheapest live provider). When registering with a preset, that list decides tool_calling unless it is given.
  • Upstream cost. OpenRouter is asked to report each call’s cost. The provider sees it in its execution record (upstream cost $…). It is never part of the signed result or the response the requester receives. InferenceUsage now tolerates unknown fields, so a later version can add a cost without breaking older requesters.
  • Model pages. A model page on openrouter.ai or huggingface.co given as an endpoint URL is refused with the preset to use instead.
  • Sharing in one step. bnw share model <MODEL> --via openrouter|huggingface publishes a router’s model, with the limits and tool support its list gives, then connects it, sets its audience, and offers it. The router’s key must be stored first. In the web client, the Share a model card (on My capabilities, and under Inference model when providing a capability) searches either router the same way and asks for the key if it is missing.
  • Registry version. The inference registry is now version 4, to hold headers and presets. Version 3 files read unchanged.

A successful provider execution appends a signed, participant-scoped InferenceExecutionClaim. It links the exact invocation, canonical capability manifest and descriptor, completed result, and encrypted output. The bounded claim distinguishes:

  • the descriptor’s public model ID, optional revision, and optional immutable artifact CID;
  • the provider’s fixed local runtime model;
  • the model string returned by the runtime, when present and valid; and
  • the provider adapter and claim method.

The initial method is provider-self-report. The signature proves which BNW provider made the statement and protects its links from modification. It does not prove that the endpoint loaded particular weights, executed the request faithfully, or returned a correct answer. bnw inference evidence <INVOCATION_ID> checks the claim against locally available invocation, descriptor, result, and output records and reports whether those links are consistent.

Claims replicate only through the requester and provider inbox scopes. BNW retains every structurally valid claim and never selects one as canonical. Multiple or contradictory claims remain visible for local policy, later third-party attestations, redundant execution, model artifact verification, or hardware-backed evidence to evaluate.