Skip to content

Model configuration and inference ​

The INFER Worker handles sampling, reasoning trajectories, reflection, deliberation and inference caching. Other Workers and application Graphs prepare context, execute tools and save results.

1. Configure a provider ​

Copy .env.example to .env and supply the actual endpoint, key and model ID:

dotenv
DITTO_SHARED_PROVIDERS=primary
DITTO_SHARED_PROVIDER_PRIMARY_KIND=openai-compatible
DITTO_SHARED_PROVIDER_PRIMARY_BASE_URL=https://your-provider.example/v1
DITTO_SHARED_PROVIDER_PRIMARY_API_KEY=YOUR_API_KEY
DITTO_SHARED_PROVIDER_PRIMARY_MODEL=YOUR_MODEL_ID
DITTO_WORKER_INFER_MODEL_PROVIDER=primary
DITTO_WORKER_INFER_MODEL=YOUR_MODEL_ID
DITTO_SHARED_SANDBOX_ALLOW_NETWORK=https://your-provider.example

primary is a registration name, not a model name. HTTP provider kinds include openai-compatible, anthropic and gemini; supported models and parameters depend on the service. The Sandbox allowlist uses URL origins, while the API base URL retains the provider's required path.

ts
import { createDitto, loadRuntimeConfig } from "@codesoul-co/ditto/runtime";
import { createInferWorker } from "@codesoul-co/ditto/worker/infer";
const config = loadRuntimeConfig(process.env);
const runtime = createDitto({ config, workers: [createInferWorker()] });

Start with node --env-file=.env app.ts. The configuration loader does not locate .env or automatically read YAML. For YAML, explicitly call loadRuntimeConfigFile("ditto.yaml",process.env). See the configuration reference for precedence, behavior settings and credentials.

2. First sample ​

ts
import { graph } from "@codesoul-co/ditto/runtime";
const answer = graph<string>("answer")
  .node("sample", "INFER.REASONING.SAMPLE", [], question => ({
    model: config.model!,
    messages: [{ role: "user", content: question }],
    generation: { maxTokens: 512 },
  }));
try {
  if (!config.model) throw new Error("Configure a model first");
  const { sample } = await runtime.run(answer, "Explain the relationship between Nodes, Workers and Graphs.");
  if (sample.status !== "success" || !sample.output) {
    throw new Error(sample.error?.message ?? "Inference failed");
  }
  console.log(sample.output.message.content);
} finally { await runtime.close(); }

Combine this code with the initialization above. It checks the inference entry point; a complete Agent also needs Context, Memory, tools and delivery. See the application entry point.

3. Convert Context into model messages ​

CONTEXT.SELECT returns ContextSelection; items are in selection.context.items. Roles loaded from messages are in item.metadata.role, not item.role. Content can be text, JSON or references and must be mapped to message types supported by the model.

For text conversations, validate system/user/assistant roles and string content before preserving the role. Treat retrieved pages and tool output as evidence or tool messages, not system instructions.

See the complete mapping in agent.ts. Provider messages and generic contract messages have different multimodal structures; a type assertion does not perform conversion.

4. Choose an inference node ​

NodeUseDownstream checks
SAMPLEOne generation or a tool action requestOuter status, message and actions
TRAJECTORYMulti-step reasoning, CoT/ToT/GoT or self-consistencyOuter success and inner completed status
REFLECTCritique, revise or iteratively improve a resultReview and revision meet application requirements
DELIBERATERank, evaluate or combine candidatesPreserve provenance and rationale; validate structure
CACHE.LOOKUP / WRITE / INVALIDATEExplicit inference cachingHit/miss, expiration and invalidation scope

These operations do not automatically create a persistent Agent. Internal inference budgets and application Loop Graph counts are different. Bound tokens, model calls, total time and external actions as appropriate.

5. Model-proposed Tool/MCP actions ​

Pass action descriptors to SAMPLE and map returned actions to an allowed call table. INFER proposes requests; it does not execute operations merely because a model emits a name. Actual actions run through INTERACTION.ACT.TOOL or an MCP Graph.

The cycle is: propose → validate name/arguments → check permissions/approval → execute → OBSERVE → update Context → infer again. Use one Loop for the cycle. See Tools and the ReAct application.

6. Providers, streaming and cache ​

ProviderRegistry.register(name,provider) accepts a ModelProvider adapter. Built-in HTTP providers use the Runtime Sandbox. Custom SDK adapters must implement network policy, cancellation and response validation.

The createInfer SDK offers streaming methods, but it is not a remote proxy for a registered Worker. Runtime invoke and ordinary Graph nodes return final results; do not assume they return the SDK's AsyncIterable. See INFER streaming.

The default inference cache is instance-local memory, independent of Redis Context and long-term Memory. Inject a shared InferCacheProvider for cross-replica reuse. Keys should include the model, input, output-affecting parameters and relevant data/prompt versions. The application controls retention of sensitive data.

7. Troubleshooting ​

ProblemCheck in order
Provider unavailableEnvironment loaded, registration names match, default model configured
Network deniedOrigin allowlist, base URL, actual proxy/SDK destination
Timeout or cancellationInference timeout, outer Loop signal, provider signal forwarding
Invalid output formatText versus structured output; validate with the application schema before writes
Unknown actionAdvertise only allowed actions; revalidate registry membership and permissions

INFER API · Providers · Configuration

Ditto · @codesoul-co/ditto · Node.js 24+