Providing from the web client
Capabilities → Provide a capability walks through everything the capability, mcp, inference, and provider commands do:
-
Describe an inference model or custom capability. This covers the name, operation, media types, size limits, privacy, and an optional JSON Schema. Publishing stores the descriptor and signs the capability manifest, exactly as
capability descriptor createandcapability publishdo.An MCP tool starts with its server instead, so each tool is published with the name, description, and input schema the server lists. Connecting the server (
provider-mcp-connect) publishes nothing: it keeps a draft connection on this node to sign in to and discover tools on. You then choose which tools to share, each becoming its own capability as withbnw share tools. Until a tool is shared, the connection appears under My capabilities as not shared yet. -
Connect a backend. Choose one of:
- a remote Streamable HTTP MCP server, with the explicit opt-in before decrypted inputs may leave the host;
- a program on this computer;
- an OpenAI-compatible or Anthropic endpoint for models, with a declared payload egress.
-
Choose the tool for a remote server. The node discovers the server’s tools, and you pin one. The pinned schema hash is always one the node saw itself. The backend forms start from the registered settings, and saving the same server again keeps the pinned tool and its sign-in.
-
Sign in for authorization-code OAuth. The node opens a loopback callback, which ignores stray browser requests, and keeps the tokens encrypted to its key, as
mcp authorizedoes. -
Secrets. You can set or remove secrets; their values are never shown again.
-
Check that the backend answers.
-
Offer the capability from this node’s peer. The node renews the offer before it lapses:
- renewal happens in the last fifth of the offer’s lifetime;
- only while the backend is ready;
- it stops once you choose “stop renewing”. Offers already signed stay valid until they expire.
-
Autopilot. Choose the mode for incoming invocations, as described above.
Two more kinds have no backend to connect:
- Search provider (
search.keyword). The descriptor is fixed: a publicsearch/querycontract. Setup is choosing who may search this node (the search policy) and granting search access to people who aren’t contacts, which signscapability:invokeon the capability. Then you offer it. The node answers from its own index, so the offer is ready and renews while search isn’t switched off. - Custom (answered by hand), for any other kind, such as
tool.translate. It has the same contract fields as an MCP tool. Invocations wait on the Invocations screen, where you grant the requester access and answer with Respond by hand. Nothing runs automatically.
Neither of them has an autopilot mode.
My capabilities lists each one with its backend, missing secrets, offer expiry, renewal, and mode.
The web client cannot register local programs by default, because a registered program runs with the node’s privileges whenever its capability is invoked. It shows the equivalent bnw mcp register command instead. To allow it, run bnw web enable --allow-local-programs and restart the node. Registering a program still asks for confirmation, naming the exact program and arguments; the program must be an absolute path. Remote servers and model endpoints can always be registered from the web client, since they only receive data the provider’s own invocations send them. Pricing terms are still attached from the CLI.
Request-level inference uses the same generic capability and invocation protocol. Create a strict descriptor, publish it as inference.model, and advertise an ordinary short-lived offer:
bnw capability descriptor create-inference inference.json \ --name "Reference text model" \ --model-id example/model-8b \ --revision v1 \ --input-modality text \ --output-modality text \ --context-tokens 8192 \ --max-output-tokens 2048
bnw capability descriptor validate inference.jsonbnw capability publish inference.model <DESCRIPTOR_ID>bnw capability offer <CAPABILITY_ID> \ --endpoint bnw.peer=<PEER_ID> \ --expires-at <UNIX_TIMESTAMP>The optional --artifact <CID> (or Model artifact CID in the web wizard) identifies an immutable model or model-manifest object; it does not force providers to distribute model weights. Model names and provider claims are descriptive rather than authoritative. The v1 inference operation is fixed to inference/generate, defaults to participant-encrypted JSON input and output, and works with the existing invoke prepare-input, request, progress, cancellation, approval, and response commands.
The reference adapter consumes strict bnw.inference-request/1 documents and produces normalized bnw.inference-response/1 documents. Provider-local configuration selects either the openai-chat or native anthropic-messages translator. The requester cannot select the adapter, endpoint, credentials, model, timeout, streaming mode, or downstream egress policy. The adapter never follows redirects, so a request and its API key go only to the registered address. A redirect fails with an error that names where it points, and registering a provider’s website (such as openai.com instead of api.openai.com) is refused with the API address to use.
For a complete validated two-Mac setup, including discovery, delegation, encrypted request preparation, execution, and result retrieval, see INFERENCE-END-TO-END.md.
For example, a provider running Ollama locally can map the descriptor’s public model claim to a fixed local runtime model. The mapping is provider-controlled and the requester cannot replace either value.
# Provider-local configuration; this file is private and is not replicated.bnw inference register <CAPABILITY_ID> \ --adapter openai-chat \ --url http://127.0.0.1:11434/v1/chat/completions \ --runtime-model gemma:2b \ --payload-egress local \ --timeout-seconds 120
bnw inference inspect <CAPABILITY_ID>bnw inference check <CAPABILITY_ID>
# Direct provider-local smoke test. This calls the model immediately and# bypasses BNW delegation, approval, encryption, and lifecycle records.cat > inference-request.json <<'JSON'{ "profile": "bnw.inference-request/1", "system": "Answer concisely.", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "Reply with one short greeting."} ] } ], "generation": {"temperature": 0.2}}JSONbnw inference call <CAPABILITY_ID> inference-request.jsonFor a remote OpenAI-compatible service, HTTPS and explicit remote payload egress are required. The secret remains in the provider environment; only its variable name is stored locally.
bnw inference register <CAPABILITY_ID> \ --adapter openai-chat \ --url https://models.example.com/v1/chat/completions \ --runtime-model provider/deployment-name \ --api-key-env MODEL_API_KEY \ --payload-egress remote
MODEL_API_KEY='provider-secret' bnw inference call \ <CAPABILITY_ID> inference-request.jsonThe native Anthropic Messages API uses the same BNW input and normalized output. The adapter changes the request shape and authentication headers behind the provider boundary:
bnw inference register <CAPABILITY_ID> \ --adapter anthropic-messages \ --url https://api.anthropic.com/v1/messages \ --runtime-model <ANTHROPIC_MODEL_ID> \ --api-key-env ANTHROPIC_API_KEY \ --payload-egress remote
ANTHROPIC_API_KEY='provider-secret' bnw inference call \ <CAPABILITY_ID> inference-request.jsonLiteLLM Proxy is an optional gateway rather than a BNW dependency. Point the openai-chat adapter at its loopback Chat Completions endpoint and keep provider credentials in LiteLLM. If LiteLLM forwards the prompt to a hosted service, declare --payload-egress remote even though BNW connects to localhost; use local only when the ultimate model execution also stays on the host.
# LiteLLM listens locally but forwards to a hosted model.bnw inference register <CAPABILITY_ID> \ --adapter openai-chat \ --url http://127.0.0.1:4000/v1/chat/completions \ --runtime-model <LITELLM_MODEL_ALIAS> \ --payload-egress remoteBNW does not inspect or control LiteLLM’s downstream logging, caching, retries, fallbacks, or routing. Those features can retain plaintext, create multiple billable upstream calls, or substitute a deployment behind an alias, so the provider must configure them consistently with the advertised capability. Existing inference registry versions migrate automatically: their former model value becomes the fixed runtime model, and version-1 entries become openai-chat. The old --model spelling remains an alias for --runtime-model, while --allow-remote-execution remains a hidden migration alias for --payload-egress remote.
The version-1 BNW request supports an optional system prompt, user and assistant messages with text content blocks, and provider-neutral maximum-output-token, temperature, top-p, and stop-sequence controls. It deliberately excludes model and streaming fields. The adapter applies the descriptor token limit, enforces input/output byte limits, disables redirects, and returns normalized text, stop reason, usage, and non-authoritative provider metadata. Rich tools, reasoning blocks, citations, multimodal data, and streaming require later versioned contracts rather than silent provider-specific fields.
To test the complete two-peer path, use the ordinary invocation flow above with that inference capability. After the request has replicated and its exact delegation or approval is valid, run this on the provider:
MODEL_API_KEY='provider-secret' \ bnw invoke execute-inference <INVOCATION_ID>
# Requester, after replication:bnw invoke status <INVOCATION_ID>bnw invoke open-payload <OUTPUT_MANIFEST_ID>bnw inference evidence <INVOCATION_ID>Execution verifies the capability/offer binding, authorization, expiry, cancellation, encrypted JSON contract, and fixed local registration before sending plaintext to the configured endpoint. It emits signed inference.accepted, inference.calling, and inference.output-prepared progress, then encrypts the normalized response for the two participants. The calling record includes only the adapter, destination origin, public claimed model, configured runtime model, and declared egress class. After successful completion, the provider also signs a participant-scoped execution claim linked to the exact invocation, capability manifest, descriptor, result, and output. It distinguishes the public descriptor claim, configured runtime model, and model name reported by the runtime. inference evidence validates and displays those links while explicitly labeling the record as a provider self-report. Every claim is retained; BNW does not choose a canonical claim or interpret agreement as correctness. A persistent provider-local claim prevents automatic replay after an ambiguous crash. None of these fields proves which weights executed or that the answer is correct.
