Inference
BNW uses a small provider-neutral JSON format at the encrypted invocation boundary. Provider-local adapters translate it to OpenAI-compatible Chat Completions, Anthropic Messages, or future model APIs. Provider endpoints, credentials, model selection, timeouts, and routing policy are never requester-controlled fields.
Request version 1
Section titled “Request version 1”{ "profile": "bnw.inference-request/1", "system": "Answer accurately and concisely.", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "What is the capital of France?"} ] } ], "generation": { "max_output_tokens": 128, "temperature": 0.2, "top_p": 0.9, "stop_sequences": ["END"] }}profile and one or more messages are required. system and generation are optional. Version 1 supports user and assistant messages containing text blocks. temperature is between 0 and 1, top_p is greater than 0 and at most 1, and at most 16 stop sequences may be supplied. A requested output-token limit cannot exceed the model descriptor’s limit; when omitted, the descriptor limit is used. The Anthropic adapter requires one of those limits because its native API requires max_tokens.
A message may also carry images, for models whose descriptor lists image among its input modalities:
{"type": "image", "media_type": "image/png", "data": "<base64>"}- Limits. Images are PNG, JPEG, GIF, or WebP, and up to 5 MiB each once decoded.
- Refusals. A provider refuses a request with images for a text-only model before accepting it, with a signed rejection.
- Adapters.
- The OpenAI-compatible adapter sends content parts with a
data:URL, but only when a message contains an image. Text-only messages are sent as plain strings, as before. - The Anthropic adapter sends native
imageblocks.
- The OpenAI-compatible adapter sends content parts with a
- Compatibility. Image blocks are an optional addition to version 1; older nodes reject requests that contain them.
A request may also offer the model tools, for providers whose offers list the inference-tools feature:
{ "profile": "bnw.inference-request/1", "messages": [ {"role": "user", "content": [{"type": "text", "text": "Price of BNW?"}]}, {"role": "assistant", "content": [ {"type": "tool_call", "id": "call-1", "name": "t1_quotes", "input": {"ticker": "BNW"}} ]}, {"role": "user", "content": [ {"type": "tool_result", "call_id": "call-1", "content": "42"} ]} ], "tools": [ {"name": "t1_quotes", "description": "Stock quotes", "input_schema": {"type": "object"}} ]}- Tools. At most 16. Each has a unique name of 1-64 letters, digits, underscores, or dashes, an optional description, and a JSON Schema object of its input up to 16 KiB.
- Calls and results. Only assistant messages carry
tool_callblocks, and each names an offered tool with a unique ID and an input object up to 16 KiB. Only user messages carrytool_resultblocks, each answering an earlier call, with up to 64 KiB of text and an optionalis_error. - Responses. A response may contain
tool_callblocks, with stop reasontool_use. The model only asks; BNW runs nothing. The requester decides what to call. - Adapters. The OpenAI-compatible adapter sends
toolsas functions, assistant calls astool_calls, and each result as atoolmessage; it readstool_callsback. The Anthropic adapter uses nativetool_useandtool_resultblocks. - Which backends take tools. A registration can say so (
bnw inference register --tool-calling true|false). Unset, the Anthropic API and OpenAI-compatible endpoints with remote egress are assumed to, and local runtimes are not, since many small local models cannot. Sharing a model in one step follows what Ollama reports. A provider refuses a request with tools for a backend that does not take them, with a signed rejection, before anything runs. - Compatibility. Tools are an optional addition to version 1, like images. Older nodes reject requests that contain them, which is why requesters send them only to offers listing
inference-tools.
The strict versioned format deliberately excludes provider model names, endpoints, credentials, streaming controls, retries, fallbacks, thinking controls, caching, and provider-specific extension objects. These can be added through later versioned contracts rather than silently changing version 1 semantics.
Response version 1
Section titled “Response version 1”{ "profile": "bnw.inference-response/1", "content": [ {"type": "text", "text": "Paris."} ], "stop_reason": "end_turn", "usage": { "input_tokens": 18, "output_tokens": 3 }, "provider": { "adapter": "anthropic-messages", "runtime_model": "provider/runtime-model-id", "request_id": "msg_example" }}Normalized stop reasons are end_turn, max_output_tokens, stop_sequence, tool_use, refusal, and other. Usage, stop sequence, provider request ID, and even the provider-reported runtime model may be absent or inaccurate. They are non-authoritative metadata signed by the BNW provider, not proof of model execution or correctness.
Version 1 normalizes returned text and tool calls, and does not expose provider-specific reasoning, citation, or cache blocks. The reference adapter is non-streaming. Richer multimodal or structured-output contracts belong in later versions.
Provider-local adapters
Section titled “Provider-local adapters”openai-chat translates the system prompt and messages to Chat Completions, inserts the fixed registered runtime model, forces stream: false, and maps the first choice into the BNW response. It is suitable for direct OpenAI-compatible services, Ollama, llama.cpp, vLLM, and an optional LiteLLM Proxy.
anthropic-messages translates the same request to the native Messages API, sends the API key using x-api-key, supplies the fixed anthropic-version, inserts the fixed registered runtime model, and normalizes the returned content, stop reason, usage, and IDs.
The public model_id in the capability descriptor and the provider’s runtime_model serve different purposes. The descriptor is the provider’s stable, discoverable claim. The local registration maps that claim to an Ollama name, hosted model identifier, deployment ID, or gateway alias. Both are provider-controlled and fixed before an invocation executes. The signed calling progress records the mapping, while the normalized result records the model string returned by the runtime; none is cryptographic proof that particular weights ran.
The local registry records payload_egress separately from the URL. local is accepted only for a loopback endpoint. remote is required whenever plaintext ultimately leaves the host, including when BNW calls a loopback LiteLLM proxy that forwards to a hosted provider. BNW cannot inspect a gateway’s downstream routing, logging, caching, retry, or fallback configuration; the provider remains responsible for configuring those behaviors consistently with its published capability.
Hosted routers: OpenRouter and Hugging Face
Section titled “Hosted routers: OpenRouter and Hugging Face”--preset openrouter and --preset huggingface set up the OpenAI-compatible adapter for a hosted router in one step. Explicit options still win.
| Preset | Endpoint | Key secret | Model names |
|---|---|---|---|
openrouter |
https://openrouter.ai/api/v1/chat/completions |
OPENROUTER_API_KEY |
vendor/model, as OpenRouter lists them |
huggingface |
https://router.huggingface.co/v1/chat/completions |
HF_TOKEN |
org/model, or org/model:provider to pin one inference provider |
bnw provider secret set OPENROUTER_API_KEYbnw inference models openrouter --search claude --toolsbnw inference register <CAPABILITY_ID> --preset openrouter --runtime-model anthropic/claude-opus-5.5- Payload egress. A preset sends prompts to the service, so egress is
remote. - Headers. OpenRouter also gets its app-attribution headers (
HTTP-Referer,X-Title).--header NAME=VALUEadds any other non-secret header, such asX-HF-Bill-To. Extra headers cannot replace the adapter’s ownAuthorization,x-api-key,anthropic-version, orAccept, and they cannot carry control characters. - Model lists.
bnw inference modelsreads the router’s public model list: context size, output limit, whether the model takes tools, and price per million tokens (for Hugging Face, the cheapest live provider). When registering with a preset, that list decidestool_callingunless it is given. - Upstream cost. OpenRouter is asked to report each call’s cost. The provider sees it in its execution record (
upstream cost $…). It is never part of the signed result or the response the requester receives.InferenceUsagenow tolerates unknown fields, so a later version can add a cost without breaking older requesters. - Model pages. A model page on
openrouter.aiorhuggingface.cogiven as an endpoint URL is refused with the preset to use instead. - Sharing in one step.
bnw share model <MODEL> --via openrouter|huggingfacepublishes a router’s model, with the limits and tool support its list gives, then connects it, sets its audience, and offers it. The router’s key must be stored first. In the web client, the Share a model card (on My capabilities, and under Inference model when providing a capability) searches either router the same way and asks for the key if it is missing. - Registry version. The inference registry is now version 4, to hold headers and presets. Version 3 files read unchanged.
Execution evidence
Section titled “Execution evidence”A successful provider execution appends a signed, participant-scoped InferenceExecutionClaim. It links the exact invocation, canonical capability manifest and descriptor, completed result, and encrypted output. The bounded claim distinguishes:
- the descriptor’s public model ID, optional revision, and optional immutable artifact CID;
- the provider’s fixed local runtime model;
- the model string returned by the runtime, when present and valid; and
- the provider adapter and claim method.
The initial method is provider-self-report. The signature proves which BNW provider made the statement and protects its links from modification. It does not prove that the endpoint loaded particular weights, executed the request faithfully, or returned a correct answer. bnw inference evidence <INVOCATION_ID> checks the claim against locally available invocation, descriptor, result, and output records and reports whether those links are consistent.
Claims replicate only through the requester and provider inbox scopes. BNW retains every structurally valid claim and never selects one as canonical. Multiple or contradictory claims remain visible for local policy, later third-party attestations, redundant execution, model artifact verification, or hardware-backed evidence to evaluate.
