v0.8.4

Architecture

Last updated

go-llm-router has three layers: the contract layer (core), the assembly layer (core/router), and the provider adapters. This page shows only the layer relationships; per-module diagrams, sequence diagrams, and state machines live in the full architecture document.

Overview

graph TB
    App[Caller] --> Router[router.New]
    Router --> Registry[newFn prefix table]
    Registry --> KeyAgents[Key-based agents]
    Registry --> OAuthAgents[OAuth agents]
    Registry --> CompatAgent[compat agent]
    KeyAgents --> Core[core contracts]
    OAuthAgents --> Core
    CompatAgent --> Core
    OAuthAgents --> OAuth[core/oauth]
    OAuth --> Keychain[go-pkg keychain]
    Core --> Stream[Stream event normalization]
    Core --> Usage[Usage normalization]
    Core --> Media[Image / audio interfaces]

Layers

Layer Packages Responsibility Dependency direction
Contract core Agent interfaces, message and usage types, reasoning levels, modes, model classification, SSE normalization, optional multimodal interfaces depends on no provider
Assembly core/router Builds an Agent from provider@model; rewrites unknown prefixes to compat depends on core and every provider
Adapters core/claude, core/openai, core/gemini, core/grok, core/deepseek, core/mistral, core/nvidia, core/ollamaCloud, core/openRouter, core/cloudflare, core/compat, core/copilot, core/openaiCodex, core/grokOauth Message conversion, wire-format differences, model listings, optional capabilities depend on core and the shared wire helpers
Shared wire helpers core/copilot/response, core/xai Responses API input/tool/output conversion; xAI Responses request body and SSE assembly depend on core
Credentials core/oauth/copilot, core/oauth/codex, core/oauth/grok Login flows, token storage and refresh depend on core and the go-pkg keychain

Adapter groups

Group Packages Wire format
OpenAI-compatible openai, deepseek, mistral, nvidia, openRouter, ollamaCloud, cloudflare, compat Chat Completions; newer OpenAI generations switch to Responses
Anthropic claude Messages API with thinking-budget conversion
Google gemini :generateContent with raw JSON Schema tools (parametersJsonSchema) and cachedContents prefix caching
xAI grok, grokOauth Responses API built by core/xai, plus the image endpoints
OAuth proxies copilot, openaiCodex Vendor-internal Responses endpoints requiring a session token

Cross-cutting principles

Principle What it means
One-way dependencies Adapters depend on core and the shared wire helpers (core/copilot/response for openai, codex, copilot; core/xai for both Grok adapters). Two adapter-to-adapter exceptions: grokOauth reuses grok.RequestImage, openaiCodex reuses the openai image helpers
Capabilities are interfaces Streaming, reasoning limits, image, STT, and TTS are optional interfaces rather than flags or config fields
Errors carry status Send returns the upstream HTTP status; stream failures wrap into *llmrouter.StreamError keeping Provider / Code / Body
Read caps live in the contract layer 64 MiB stream body, 8 KiB error body, 512-byte error frame, 64 KiB JSON body, all defined once in core
Shared HTTP client llmrouter.NewHTTPClient gives a 10-minute timeout; codex and grok-oauth build their own client with a 15-second response-header timeout

Further reading

中文