Connect model tools to the paid console
Point the coding harness you already run at the paid console’s model endpoint, https://api.muniment.ai. The console holds provider keys and budgets, so your machine holds no provider keys.
Access
Access path
- The paid console is not yet open to the public.
- A console owner opens your user record in Effective permissions and selects Regenerate virtual key. The first confirmation reads Issue virtual key. The owner issues your virtual key, not a provider key.
- The short-lived agent bearer comes from the agent authentication flow.
Key safety
- The virtual key and short-lived agent bearer are identity credentials.
- Keep both credentials out of source control.
- Keep both credentials out of command history.
- Keep both credentials out of support messages.
- Keep both credentials out of application logs.
Choose a client recipe
Choose the client you already run to see its verified recipe for the paid console’s model endpoint.
| Client | Wire protocol | Credential |
|---|---|---|
| Claude Code | Anthropic Messages | virtual key |
| Codex CLI | OpenAI Responses | virtual key |
| aider | OpenAI-compatible chat completions | virtual key |
| OpenCode | OpenAI-compatible chat completions | virtual key |
| OpenAI Agents SDK for Python | OpenAI-compatible chat completions | virtual key |
| OpenAI Agents SDK for TypeScript | OpenAI-compatible chat completions | virtual key |
| CrewAI 1.15.2 | OpenAI-compatible chat completions | short-lived agent bearer |
| PydanticAI 2.9.1 | OpenAI-compatible chat completions | short-lived agent bearer |
The paid console’s model endpoint accepts three wire protocols:
- Anthropic Messages at
/v1/messages - OpenAI Responses at
/v1/responses - OpenAI-compatible chat completions at
/v1/chat/completions
The protocol does not change access. The console classifies each prompt and selects an approved deployment whose registry tier matches that class. The model name from a virtual-key client is a preference among matching deployments, not a selector. muniment may serve a different model. The console applies the same virtual key budget across tools. Successful calls produce usage receipts for the organization and user that the virtual key identifies.
Claude Code
Claude Code uses the Anthropic Messages protocol. Set these variables in the shell that launches claude:
export ANTHROPIC_BASE_URL="https://api.muniment.ai"
export ANTHROPIC_AUTH_TOKEN="<muniment-virtual-key>"
unset ANTHROPIC_API_KEY
claude --model "<model-preference>"Claude Code supplies the required anthropic-version header. Feature-specific anthropic-beta headers are passed through unchanged.
For fleet-managed installations, deploy the following managed-settings.json through MDM or the platform’s system-level managed-settings location:
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.muniment.ai"
}
}Inject ANTHROPIC_AUTH_TOKEN separately for each identity. Do not put a shared virtual key in managed settings.
Codex CLI
Export the key, then add a provider to ~/.codex/config.toml:
export MUNIMENT_API_KEY="<muniment-virtual-key>"model_provider = "muniment"
model = "<model-preference>"
[model_providers.muniment]
name = "Muniment"
base_url = "https://api.muniment.ai/v1"
env_key = "MUNIMENT_API_KEY"
wire_api = "responses"Codex sends OpenAI Responses requests. Keep wire_api = "responses". Chat completions is not a Codex CLI wire mode.
aider
aider uses OpenAI-compatible chat completions for this connection:
export OPENAI_API_BASE="https://api.muniment.ai/v1"
export OPENAI_API_KEY="<muniment-virtual-key>"
aider --model "openai/<model-preference>"If the model is not in aider’s built-in metadata, its model warning does not grant or deny access. The muniment server-side entitlement remains authoritative.
OpenCode
Export the key:
export MUNIMENT_API_KEY="<muniment-virtual-key>"Add a custom OpenAI-compatible provider to opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"providers": {
"muniment": {
"package": "aisdk:@ai-sdk/openai-compatible",
"name": "Muniment",
"settings": {
"baseURL": "https://api.muniment.ai/v1",
"apiKey": "{env:MUNIMENT_API_KEY}"
},
"models": {
"<model-preference>": {
"name": "<model-preference>"
}
}
}
}
}Select muniment/<model-preference> in /models. This sample uses the OpenAI-compatible chat completions wire protocol. For an OpenCode model that specifically requires Responses, use "package": "aisdk:@ai-sdk/openai" with the same base URL and key.
OpenAI Agents SDK for Python
Install openai-agents, export an identity’s virtual key and model preference, then run this complete example:
python -m pip install openai-agents
export MUNIMENT_API_KEY="<muniment-virtual-key>"
export MUNIMENT_MODEL="<model-preference>"import os
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled
from openai import AsyncOpenAI
key = os.environ["MUNIMENT_API_KEY"]
model_name = os.environ["MUNIMENT_MODEL"]
# Muniment governs model calls, but does not accept OpenAI trace exports.
set_tracing_disabled(True)
client = AsyncOpenAI(
api_key=key,
base_url="https://api.muniment.ai/v1",
)
agent = Agent(
name="Muniment assistant",
instructions="Answer concisely.",
model=OpenAIChatCompletionsModel(
model=model_name,
openai_client=client,
),
)
result = Runner.run_sync(agent, "Name one benefit of least-privilege access.")
print(result.final_output)OpenAI Agents SDK for TypeScript
Create a project and install the SDKs and TypeScript runner:
npm init -y
npm install @openai/agents openai tsxExport the same two variables shown in the Python example and save this complete example as muniment-agent.ts:
import {
Agent,
run,
setDefaultOpenAIClient,
setOpenAIAPI,
setTracingDisabled,
} from "@openai/agents";
import OpenAI from "openai";
async function main() {
const apiKey = process.env.MUNIMENT_API_KEY;
const model = process.env.MUNIMENT_MODEL;
if (!apiKey || !model) {
throw new Error("Set MUNIMENT_API_KEY and MUNIMENT_MODEL");
}
setDefaultOpenAIClient(new OpenAI({
apiKey,
baseURL: "https://api.muniment.ai/v1",
}));
setOpenAIAPI("chat_completions");
setTracingDisabled(true);
const agent = new Agent({
name: "Muniment assistant",
instructions: "Answer concisely.",
model,
});
const result = await run(
agent,
"Name one benefit of least-privilege access.",
);
console.log(result.finalOutput);
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});Run it with the installed TypeScript runner:
npx tsx muniment-agent.tsThese examples deliberately select chat completions, the OpenAI-compatible agent-framework path verified by the gateway integration suite. Features that require OpenAI-hosted services, including OpenAI trace export, are not routed through muniment.
CrewAI 1.15.2
The short-lived agent bearer comes from the agent authentication flow.
Unlike the virtual-key model preference, <entitled-model> names a model the agent’s organization entitles. The gateway refuses a model entitled only to another organization.
Install the exact version exercised by the gateway integration suite:
python -m pip install "crewai==1.15.2"
export MUNIMENT_AGENT_BEARER="<short-lived-muniment-agent-bearer>"
export MUNIMENT_MODEL="<entitled-model>"Use CrewAI’s public LLM configuration with the muniment URL and agent bearer. The explicit provider="openai" selects CrewAI’s OpenAI-compatible chat client without rewriting or restricting muniment’s public model name.
import os
from crewai import LLM
llm = LLM(
model=os.environ["MUNIMENT_MODEL"],
provider="openai",
api_key=os.environ["MUNIMENT_AGENT_BEARER"],
base_url="https://api.muniment.ai/v1",
max_tokens=64,
)
print(llm.call("Name one benefit of least-privilege access."))Do not substitute an OpenAI or other provider key. Version 1.15.2 is verified to honor the configured muniment base URL without OpenAI-host key validation.
PydanticAI 2.9.1
Install the exact verified version and inject the same kind of short-lived agent bearer:
python -m pip install "pydantic-ai==2.9.1"
export MUNIMENT_AGENT_BEARER="<short-lived-muniment-agent-bearer>"
export MUNIMENT_MODEL="<entitled-model>"import os
from pydantic_ai import Agent
from pydantic_ai.models.openai import OpenAIChatModel
from pydantic_ai.providers.openai import OpenAIProvider
model = OpenAIChatModel(
os.environ["MUNIMENT_MODEL"],
provider=OpenAIProvider(
base_url="https://api.muniment.ai/v1",
api_key=os.environ["MUNIMENT_AGENT_BEARER"],
),
)
agent = Agent(model)
print(agent.run_sync("Name one benefit of least-privilege access.").output)These two configurations are connection proofs: the suite verifies one non-streaming chat-completions result per framework, tenant-scoped receipt provenance, provider-boundary redaction, and rejection of a model entitled only to another organization. Framework tools, capabilities, streaming, and framework-managed tracing are outside this verification scope.
Other OpenAI-compatible tools
Configure the tool’s OpenAI base URL as https://api.muniment.ai/v1, its API key as your muniment virtual key, and its model as a model preference. The tool must use either /v1/chat/completions or /v1/responses. Changing the URL to a provider endpoint bypasses muniment and does not produce a muniment usage receipt.
Troubleshooting
400withinvalid_model_request: the request body or model field is invalid. Correct the request body or model preference, then try again.401: the virtual key is missing, invalid, expired, revoked, or blocked. Ask a console owner to check the virtual key. If it expired, ask the owner to use Regenerate virtual key in Effective permissions. The paid console is not yet open to the public. If you need public access, join the waitlist for updates.403withagent_auth_forbidden: the agent-bearer path refused the authenticated agent. This status does not apply to the virtual-key path. Ask a console owner to check the agent’s permissions before you try again.502: the gateway or configured model provider is temporarily unavailable. Try again later. If the error continues, report it to a console owner.503withmodel_gateway_admission_unavailable: your organization has no approved model for the classified tier or the configured default tier. Try again after a console owner approves a matching model.
Never replace the virtual key with a provider key. The console holds provider keys, not your machine.
Verified client scope
The gateway integration suite exercises Anthropic Messages, OpenAI Responses, and OpenAI-compatible chat completions with real virtual-key model scope, budget enforcement, provider routing, and usage receipts. It also verifies forwarding of Claude Code’s anthropic-version and anthropic-beta headers.
The Python and TypeScript agent examples above use the agent SDKs’ documented custom OpenAI client and chat-completions selection against that verified protocol. The TypeScript example’s documented tsx execution path was also verified with the current packages.
CrewAI 1.15.2 and PydanticAI 2.9.1 are additionally exercised as real pinned clients with short-lived agent bearers, including a cross-tenant model denial before provider dispatch.
This page makes support claims only for the clients and configurations shown above. Claude Desktop, Claude’s VS Code extension, Codex desktop, and subscription-auth interactions have not been verified against the shipped gateway and are intentionally not described as supported configurations.
Operate the provider vault
Operators can configure the provider vault, rotate its local key, and migrate it to AWS KMS.