Provider API
src/worker/infer/providers/ contains one registry, one model interface and one HTTP/SSE transport. OpenAI-compatible, Anthropic and Gemini each have a protocol file, with no vendor SDK dependencies. Cache backends live in infer/cache/provider.ts.
Interface and registration
import type { SampleInput, SampleOutput } from "@codesoul-co/ditto/worker/infer";
export type ModelStreamEvent =
| { type: "text_delta"; delta: string }
| { type: "result"; output: SampleOutput };
export interface ModelProvider {
invoke(input: SampleInput, options: { signal: AbortSignal }): Promise<SampleOutput>;
stream?(input: SampleInput, options: { signal: AbortSignal }): AsyncIterable<ModelStreamEvent>;
}invoke is required; stream is optional. Streams yield public text deltas, then exactly one complete result with message, actions and usage, then end. Providers return SampleOutput; INFER adds the NodeResult envelope. Custom providers should honor signal. The SDK stops waiting for uncooperative calls but cannot forcibly terminate their work.
import { ProviderRegistry } from "@codesoul-co/ditto/worker/infer/providers";
const providers = new ProviderRegistry({
fixture: { invoke: async input => ({
message: { role: "assistant", content: `Received ${input.messages.length} messages` },
finishReason: "stop",
}) },
});
const remove = providers.register("other", anotherProvider);
const provider = providers.get("fixture");
remove();| Method | Semantics |
|---|---|
new ProviderRegistry(providers?) | Optional readonly name → ModelProvider map; empty by default |
register(name, provider): () => boolean | Reject empty/duplicate names; return a function removing only this registration |
get(name?): ModelProvider | Exact lookup when named; omitted name selects a sole entry; missing/ambiguous selection throws PROVIDER_NOT_FOUND |
createInfer({ providers }) and createInferWorker({ providers }) accept a registry or name map. Workers default to ctx.services.providers; createInfer({ runtime }) shares the same registry. Selection: input.model.provider → defaultProvider → Runtime config.model.provider → sole registry entry. Explicit providers replace, rather than merge with, the Runtime registry. CACHE works without model providers.
HTTP providers
import { createHttpProvider } from "@codesoul-co/ditto/worker/infer/providers";
const provider = createHttpProvider({
kind: "anthropic", baseUrl: "https://api.anthropic.com/v1",
...(process.env.ANTHROPIC_API_KEY ? { apiKey: process.env.ANTHROPIC_API_KEY } : {}),
sandbox: runtime.services.sandbox, timeoutMs: 30_000,
});
runtime.services.providers.register("claude", provider);| Option | Meaning |
|---|---|
| kind | "openai-compatible" | "anthropic" | "gemini" |
| baseUrl | Required API base, not method URL; HTTP(S), without credentials/query/hash |
| apiKey? | Provider configuration only; optional for local endpoints |
| sandbox | Required Sandbox implementing assert("network", origin), checked on every request |
| timeoutMs? | Default 30,000; integer 1–2^31−1; covers connection/body reads and combines with caller signal |
| model? | Provider model ID for applications/live runner; Node input still supplies model explicitly |
| maxTokensField? | OpenAI-compatible token-limit field: max_completion_tokens (default) or max_tokens for DeepSeek/Zhipu |
| providerOptions? | Provider body defaults, e.g. thinking; input.model.providerOptions overrides defaults; generation takes precedence |
| fetch? | Optional injected fetch; defaults to globalThis.fetch |
If supplied, input.model.endpoint must match baseUrl. Redirects are rejected. Non-2xx responses produce PROVIDER_HTTP_ERROR with status, without echoing response bodies. No automatic retries, vendor fallback or model catalog.
| Protocol | Method path | Authentication/protocol headers | Output token limit |
|---|---|---|---|
| OpenAI-compatible | /chat/completions | Authorization: Bearer … | max_completion_tokens or configured max_tokens |
| Anthropic | /messages | x-api-key, anthropic-version: 2023-06-01 | max_tokens; defaults to 4096 |
| Gemini | /models/{model}:generateContent; streaming :streamGenerateContent?alt=sse | x-goog-api-key | generationConfig.maxOutputTokens |
Generation temperature/topP/topK/stop map to protocol fields. Seed is forwarded for OpenAI-compatible/Gemini and rejected for Anthropic. Individual model support varies. String messages are portable; array content is vendor-native, without automatic cross-protocol multimodal conversion.
model.providerOptions adds body fields. The adapter controls model/messages/tools/stream/candidate count; explicit generation takes precedence. OpenAI fixes n=1; Gemini fixes candidateCount=1 and merges providerOptions.generationConfig with generation. Reflection/deliberation use JSON prompts and validation; provider-specific schema modes may be configured through providerOptions.
Tool history and streaming
SAMPLE returns actionRequests without execution. Use the ReAct Runtime flow or application Graphs to execute them. Caller declarations control names and routing targets; target Workers validate business arguments.
History uses assistant.metadata.actionRequests and tool.metadata.actionRequestId/name. Tool content is text or JSON text. These map to OpenAI tool_calls/tool_call_id, Anthropic tool_use/tool_result, and Gemini functionCall/functionResponse.
Keep the returned message.metadata intact: Anthropic contentBlocks preserve signed blocks; Gemini parts preserve thoughtSignature. Replay native blocks within the same vendor; do not assume they survive cross-vendor switching. Native Gemini function IDs are replayed when supplied; otherwise local IDs associate observations without inventing wire IDs.
A shared decoder handles split UTF-8/CRLF SSE, caps individual frames at 1 Mi JavaScript characters, and releases readers on exit. OpenAI requires [DONE], Anthropic requires message_stop and closed content blocks, and Gemini requires finishReason. Missing termination produces INCOMPLETE_MODEL_OUTPUT. Only public text becomes text_delta; thinking/signature blocks remain opaque replay data.
Usage comes from provider counters; missing counters are not fabricated. Anthropic input includes cache creation/read tokens. Gemini outputTokens includes thoughtsTokenCount; reasoningTokens is also reported but not counted twice in totalTokens. Token budgets require totalTokens or complete input/output accounting per call, otherwise USAGE_UNAVAILABLE. Input token cost is unknown beforehand, so this is not a hard billing cap.
Environment and multiple providers
import { createDitto, createInferWorker, loadRuntimeConfigFile } from "@codesoul-co/ditto";
const config = loadRuntimeConfigFile("ditto.yaml", process.env);
const runtime = createDitto({ config, workers: [createInferWorker()] });
if (!config.model) throw new Error("Configure DITTO_WORKER_INFER_MODEL_PROVIDER and DITTO_WORKER_INFER_MODEL");
const response = await runtime.invoke("INFER.REASONING.SAMPLE", {
model: config.model, messages: [{ role: "user", content: "Hello" }],
});
await runtime.close();Load the example configuration explicitly: DITTO_SHARED_PROVIDERS=deepseek,openai,glm,claude,gemini,local, with DITTO_SHARED_PROVIDER_<NAME>_KIND/BASE_URL/API_KEY/MODEL for each entry. Names match [a-z][a-z0-9_]* and must be unique. Default base URLs by kind are https://api.openai.com/v1, https://api.anthropic.com/v1, and https://generativelanguage.googleapis.com/v1beta. Allow the corresponding origins through DITTO_SHARED_SANDBOX_ALLOW_NETWORK.
DITTO_WORKER_INFER_MODEL_PROVIDER and DITTO_WORKER_INFER_MODEL must be configured together; callers supply the model field on each request. Runtime never implicitly loads environment/files. Supplying createDitto({ providers }) skips provider construction from config.providers.
Provider behavior is configured in YAML as shared.providers.<name>.options and maxTokensField (OpenAI-compatible only). See the shared configuration API for defaults and overrides. Keep real configuration in the ignored root .env. Every Worker shares ctx.services.config/providers/sandbox; unused storage options are not invented. Reasoning models may count internal reasoning against maxTokens, so an exhausted budget produces length/partial rather than a complete success. Compatible tool messages also retain original reasoning_content when supplied, for replay only, never text_delta.
Protocol references: Anthropic streaming, Gemini generation, Gemini function calling.
Examples for each API
Complete source: examples/infer.ts. The functions below share its imports; importing the file executes no examples. Applications supply database, model, or MCP resources. Choose the function you need; writes, deletes, and model calls perform real operations when invoked.
import { createDitto, loadRuntimeConfigFile } from "@codesoul-co/ditto";
import {
createInfer, createInferWorker, InMemoryInferCache, inferSampleNode,
type InferClient, type ModelConfig, type TrajectoryInput, type ReflectInput,
type DeliberateInput, type TrajectoryStrategy, type InferCacheProvider, type ModelProvider, type SampleInput,
} from "@codesoul-co/ditto/worker/infer";
import { ProviderRegistry, createHttpProvider, type HttpProviderOptions } from "@codesoul-co/ditto/worker/infer/providers";ProviderRegistry.register / get / unregister and invoke
Shows registration, resolution, raw invocation, and removal. ModelProvider.invoke returns SampleOutput without NodeResult. Removal neither closes resources nor cancels in-flight calls.
export async function providerRegistryApis(provider: ModelProvider, input: SampleInput) {
const providers = new ProviderRegistry();
const unregister = providers.register("primary", provider);
try {
const selected = providers.get("primary");
return await selected.invoke(input, { signal: AbortSignal.timeout(5_000) });
} finally { unregister(); }
}ModelProvider.stream: raw stream
stream is optional. Raw streams contain text_delta and one complete result, without SDK start/step events. Falling back to invoke does not provide token streaming.
export async function providerStream(provider: ModelProvider, input: SampleInput) {
const signal = AbortSignal.timeout(5_000);
if (!provider.stream) return provider.invoke(input, { signal });
for await (const event of provider.stream(input, { signal })) {
if (event.type === "text_delta") process.stdout.write(event.delta);
if (event.type === "result") return event.output;
}
throw new Error("Provider stream ended without a result");
}createHttpProvider: explicit construction
Options are documented above. sandbox must allow the baseUrl origin. Construction creates an adapter without connecting; invoke/stream performs requests.
export function httpModelProvider(options: HttpProviderOptions) {
return createHttpProvider(options);
}
// Example options: { kind: "openai-compatible", baseUrl: "https://api.openai.com/v1",
// apiKey: process.env.DITTO_SHARED_PROVIDER_OPENAI_API_KEY, sandbox: runtime.services.sandbox }
// Omit apiKey entirely when the endpoint has no authentication.