Project layout¶
Package structure¶
sira/ # the Python package
├── main.py # Typer CLI: tailor, re-tailor, resume, runs, setup. Scraping happens HERE.
├── paths.py # per-user data directory (SIRA_DATA_DIR)
├── durability.py # DBOS runtime: durable_runtime(), checkpoint DB URL
├── __main__.py # `python -m sira` entry point
├── workflows/
│ ├── __init__.py # ResumeTailorWorkflow — the 6-stage pipeline, as the sira.tailor DBOS workflow
│ ├── agents.py # every agent, plus model/quality-gate machinery
│ ├── continuation.py # sira resume: resume interrupted runs, fork failed ones
│ └── skill_matching.py # match_skills: pre-pass → skill judge → fallback
├── models/
│ ├── agents/
│ │ ├── output.py # the typed contracts between stages
│ │ └── deps.py # agent dependency types
│ └── workflow.py # ResumeTailorResult
├── memory/
│ ├── models.py # domain models (ResolvedOriginalResume, …)
│ ├── parser.py # PydanticAIResumeParser (adapter)
│ ├── repository.py # abstract ResumeMemoryRepository interface
│ ├── sqlite_repository.py # the SQLite implementation
│ └── service.py # ResumeMemoryService — orchestration + caching
├── reporting/
│ ├── base.py # ProgressReporter protocol, use_reporter(), NullReporter
│ ├── dashboard.py # LiveDashboard (Rich); degrades in a non-TTY
│ └── verbose.py # VerboseReporter (--verbose)
├── tools/
│ ├── job_scraper.py # fetch_job_markdown: Playwright → Markdown → assert_quality → injection scan
│ ├── job_scraper_helpers.py # HTML→Markdown, placeholder + prompt-injection detection, cleanup
│ └── injection_guard.py # optional local classifier (sira[guard] extra) + consent flow
├── rendering/
│ ├── __init__.py # render_resume(cv, dir, base_name, style) → .md/.pdf/.docx
│ ├── errors.py # RenderError
│ ├── templates.py # TemplateSpec: modern, classic, compact
│ ├── inline.py # inline markdown subset (links, bold, italic, code)
│ ├── html.py + resume.html.j2 # CV → HTML (Jinja2)
│ ├── css.py # TemplateSpec → CSS for the PDF
│ ├── pdf.py # HTML + CSS → PDF (PyMuPDF Story)
│ ├── docx.py # CV + TemplateSpec → DOCX (python-docx)
│ └── markdown.py # CV → Markdown
└── utils/ # no model calls live in here
├── cv_diff.py # CVDiff + GapAnalysis + match score, pure Python
├── skill_matching.py # CV text rendering + literal skill pre-pass
├── skill_cleanup.py # collapse duplicate skill variants (pure Python)
├── markdown_writer.py # generate_report_markdown
├── resume_converter.py # DOCX/PDF → Markdown (markitdown)
└── validate_inputs.py # deprecated, unused by the CLI
tests/
├── conftest.py # blocks real model calls; resets global model state
├── factories.py # test data builders
├── memory/ # memory layer
├── reporting/ # reporters
├── workflows/ # pipeline, loop config, model tiers, reporter events
└── test_*.py # CLI, diff, scraper, quality gate, converters, smoke
output/ # default output directory (gitignored)
docs/ # this site
mkdocs.yml # site configuration and navigation
Where the surprises are¶
Reading the code top-down does not tell you the execution order, because the pipeline is not all inside the workflow class.
sequenceDiagram
autonumber
participant U as You
participant CLI as main.py
participant MEM as ResumeMemoryService
participant FETCH as fetch_job_markdown()
participant SCR as job_scraper_agent
participant WF as ResumeTailorWorkflow (DBOS)
U->>CLI: sira tailor URL RESUME
CLI->>CLI: convert DOCX/PDF to Markdown
CLI->>CLI: apply_model_override(--model)
CLI->>MEM: resolve original resume (cache by content hash)
CLI->>FETCH: fetch the posting (Playwright, no model)
FETCH-->>CLI: RawScrape (Markdown + injection indicators)
CLI->>SCR: strip site chrome
SCR-->>CLI: cleaned posting Markdown (str)
CLI->>WF: run(resume text, posting Markdown)
WF-->>CLI: ResumeTailorResult
CLI->>MEM: store tailored resume + audit + posting
CLI->>U: write .md/.pdf/.docx + report, print report
Three things that catch people out:
- Scraping happens in
main.py, before the workflow. So does resolving the original resume from memory and applying--model. The workflow receives text, not a URL. --modelmutates module-level globals inworkflows/agents.py, and is applied before the scraper for exactly that reason.- Gap analysis, match score and verdict are computed in pure Python in
utils/cv_diff.pyfrom the skill matcher's per-skill verdicts. The report agent writes prose around numbers it did not choose.
Progress reporting¶
sira/reporting/ is a small, context-local abstraction for showing progress. It exists
so that agent code never has to know whether anyone is watching.
ProgressReporterinreporting/base.pyis aProtocol— a structural interface, so any object with the right methods qualifies; no base class to inherit from.- The active reporter lives in a
contextvars.ContextVar. Get it withget_active_reporter(), install one for the current async context with theuse_reporter(reporter)context manager. run_agent()emits stage start/finish, retry, quality-score, and token-streaming events to whichever reporter is active.
| Implementation | Used when |
|---|---|
NullReporter |
tests, and any time no reporter is installed — does nothing |
LiveDashboard |
the default. A live Rich panel in a terminal; plain line-by-line logging in a non-TTY such as CI or a pipe |
VerboseReporter |
--verbose. Streams thinking and output tokens straight to stdout, no live panel |
Reporter calls are best-effort. _safe_report() swallows exceptions from a reporter so
that a display bug can never abort a pipeline run.
Known limitation
The workflow still writes some progress lines with bare print(). In an
interactive terminal those can interleave with the LiveDashboard live panel and
garble the display — cosmetic only. Routing them through the reporter is the fix.
Speed levers¶
Four independent mechanisms reduce end-to-end latency. --fast turns on all four.
- Parallel parse and analyse. On a cold cache, stages 1 and 2 run concurrently as
two DBOS child workflows (
DBOS.start_workflow_async), each owning its own step sequence. Notasyncio.gather— interleaving two agent runs inside one DBOS workflow breaks checkpoint replay. - Advisory quality gate. One scoring pass, and a retry only below the threshold — not a loop until perfect. Parser and Analyst are not gated at all.
- Trimmed loops. Defaults are 2 write attempts × 1 review iteration, adjustable
with
--write-attemptsand--review-iterations. - Per-agent model tiers.
set_agent_models(fast=…, strong=…)andresolve_model(label)put mechanical stages on a cheaper model.
Which file to change¶
| You want to | Edit |
|---|---|
| Add or change a CLI flag | sira/main.py (both commands) |
| Change what an agent is told to do | sira/workflows/agents.py (system prompts) |
| Change the shape of a stage's output | sira/models/agents/output.py |
| Change the loop, retries, or fallbacks | sira/workflows/__init__.py |
| Change what is stored, or the cache rules | sira/memory/service.py |
| Change how progress is displayed | sira/reporting/ |
| Change how the resume is rendered, or add a style | sira/rendering/ |
| Change how the report file is written | sira/utils/markdown_writer.py |
| Change scraping or HTML extraction | sira/tools/job_scraper.py (fetch) + sira/tools/job_scraper_helpers.py (parsing) + the job_scraper_agent prompt in sira/workflows/agents.py |
| Change how runs are resumed or forked | sira/workflows/continuation.py, sira/durability.py |
Most changes land in workflows/agents.py. It is the largest and most central file.
Files that are not what they look like¶
sira/utils/validate_inputs.pyand theruntarget in theMakefileare deprecated and broken. Useuv run sira ….cover_letter_writer_agentexists inagents.pybut is not wired into the workflow.job_scraper_agentis the scraper the CLI uses; the old tool-basedscraper_agentwas removed.ScrapedJobPostinginmodels/agents/output.pyis legacy; the live path uses theRawScrapedataclass fromtools/job_scraper.pyand passes plain Markdown into the workflow._parser_qsand_analyst_qsinagents.pyare read by fallback branches but never written — the parser and analyst gates were removed for speed.