Models and providers¶
Sira is built on PydanticAI, which speaks to many LLM
providers behind one interface. You pick one with --model, using the format
<provider>:<model>.
The default is openai:gpt-5-mini.
uv run sira tailor <JOB_URL> <RESUME_PATH> --model anthropic:claude-sonnet-4-5
Supported providers¶
| Provider | Prefix | Example --model |
Environment variable |
|---|---|---|---|
| OpenAI | openai: |
openai:gpt-4o-mini |
OPENAI_API_KEY |
| Anthropic | anthropic: |
anthropic:claude-sonnet-4-5 |
ANTHROPIC_API_KEY |
| Google Gemini | google: |
google:gemini-3-pro-preview |
GOOGLE_API_KEY |
| Google Cloud (Vertex AI) | google-cloud: |
google-cloud:gemini-3-flash-preview |
GOOGLE_API_KEY |
| Groq | groq: |
groq:llama-3.3-70b-versatile |
GROQ_API_KEY |
| Mistral | mistral: |
mistral:mistral-large-latest |
MISTRAL_API_KEY |
| xAI | xai: |
xai:grok-3-mini |
XAI_API_KEY — needs the xai extra (see note below) |
| Cohere | cohere: |
cohere:command-r-plus |
CO_API_KEY |
| DeepSeek | deepseek: |
deepseek:deepseek-chat |
DEEPSEEK_API_KEY |
| OpenRouter | openrouter: |
openrouter:openai/gpt-4o |
OPENROUTER_API_KEY |
| Ollama (local) | ollama: |
ollama:llama3 |
OLLAMA_BASE_URL |
| GitHub Models | github: |
github:xai/grok-3-mini |
GITHUB_API_KEY — retired by GitHub on 2026-07-30; removed in pydantic-ai v3 |
| Cerebras | cerebras: |
cerebras:llama3.1-8b |
CEREBRAS_API_KEY |
| AWS Bedrock | bedrock: |
bedrock:anthropic.claude-sonnet-4-5 |
AWS credentials |
Every provider above works out of the box except the retired GitHub Models and
xAI: pydantic-ai 2.x uses the native xai-sdk, which cannot be installed together with Sira's dev tools, so
it is an opt-in extra — install with uv sync --extra xai --no-dev instead of
plain uv sync.
How a model reaches an agent¶
Every agent is created once, at import time, with a shared default model object. The
--model value replaces that at run time, per call. Two tiers exist so that
mechanical stages can run on a cheaper model than the ones doing the real writing.
flowchart TD
CLI["--model / --fast on the command line"] --> AMO["apply_model_override()"]
AMO --> G["Module globals:<br/>MODEL_NAME, FAST_MODEL, STRONG_MODEL"]
G --> RM{"resolve_model(agent_label)"}
RM -->|"label is Parser, Analyst, Reviewer,<br/>Quality Gate, Skill Matcher, Scraper"| FAST["FAST tier"]
RM -->|"label is Writer, Auditor,<br/>Report, Cover Letter Writer"| STRONG["STRONG tier"]
RM -->|"no override configured"| DEF["None -> the agent's own<br/>import-time default model"]
FAST --> RUN["agent.run(model=...)"]
STRONG --> RUN
DEF --> RUN
Two consequences worth knowing:
- The override is applied before the scraper runs, not inside the workflow. The
scraper and the cached resume parser run outside the pipeline, and they honour
--modeltoo. - Without
--fastor an explicit tier configuration, both tiers point at the same model, so--modelsimply applies everywhere.
Which stages use which tier¶
| Tier | Stages |
|---|---|
| fast | Resume Parser, Job Analyst, Reviewer, Quality Gate, Skill Matcher, Job Scraper |
| strong | CV Writer (initial and refine), Auditor, Report, Cover Letter Writer |
The mapping is _AGENT_TIERS in sira/workflows/agents.py; an unknown label falls back
to the strong tier.
--fast sets the fast tier to openai:gpt-5-nano and the strong tier to openai:gpt-5-mini,
or to your --model value if you passed one.
--fast always uses an OpenAI model for the fast tier
Combining --fast with, say, --model anthropic:claude-sonnet-4-5 puts Anthropic
on the strong tier but leaves openai:gpt-5-nano on the fast tier — so you still
need an OPENAI_API_KEY. If you want a single provider end to end, use --model
on its own and tune --write-attempts / --review-iterations by hand.
Running a local model with Ollama¶
Ollama runs models on your own machine. PydanticAI talks to it through its OpenAI-compatible endpoint, and requires you to say where that endpoint is — there is no default:
ollama pull llama3
export OLLAMA_BASE_URL=http://localhost:11434/v1
uv run sira tailor <JOB_URL> <RESUME_PATH> --model ollama:llama3
Ollama's cloud models route through the same local daemon. Sign in first with
ollama signin:
export OLLAMA_BASE_URL=http://localhost:11434/v1
uv run sira tailor <JOB_URL> <RESUME_PATH> --model 'ollama:kimi-k2.6:cloud'
You do not need an OpenAI key for this
Agents are deliberately constructed without touching provider credentials at
import time. If you run with --model ollama:… and no OPENAI_API_KEY, import
still succeeds and no OpenAI call is ever made.
Structured output is the hard part for local models
Every stage returns a typed object, not free text. Smaller local models are worse at producing valid structured output, so expect more retries and occasional fallbacks. The quality gate and the retry loops absorb some of that, but a cloud provider is more reliable.
OpenAI-compatible providers¶
Many services expose an OpenAI-compatible API. PydanticAI reaches them through the
openai: prefix plus provider-specific environment variables. See the
PydanticAI OpenAI docs for
Together AI,
Perplexity,
Fireworks AI, and
Azure AI Foundry.
Controlling cost¶
Roughly in order of impact:
--no-quality-gate— removes the scoring call that follows each gated agent.--fast— a cheaper tier for mechanical stages plus a lower gate threshold (5). It leaves the loops at their defaults (2 write attempts × 1 review iteration).--write-attempts 1 --review-iterations 0— a single pass with no refinement.- A cheaper
--model.
A hard ceiling also exists in code: USAGE_LIMITS = UsageLimits(request_limit=1000)
caps the number of model requests in a single run.