For engineering teams
The state of the codebase, every sprint.
Watchdog is a survey on your calendar, not a gate in your pipeline. Every sprint it reads the whole product — every service, every language — and tells you what moved: the CAI, the trend on each lens, the slop-vs-brilliant composition of what you merged, architecture and security drift, and a ranked fix list your coding agent can act on. The next survey shows whether it worked.
Self-serve · no credit card · nothing installed · read-only · the first full report on any repo is free.
Every sprint
a full survey on your team's cadence — never a gate in the pipeline
Daily
security watch — CVEs, secrets, regressions
10
lenses — 5 always on, 5 light up with your architecture
100%
reproducible — same commit, same rubric, same number
€0
first scan on every repo
What a team actually gets
Four arguments you can't make today.
An answer to "is it getting better?"
One number, the same rubric every sprint. The tech lead's opinion and the new hire's opinion stop being the whole evidence base.
A case for the refactor
"The code is bad" loses to a roadmap item every time. A cohesion lens that has dropped four surveys running, with file:line evidence, does not. And when the work is done, the next survey shows the movement — so the next refactor is an easier conversation.
An architecture diagram that's true
Redrawn from the actual dependencies on every survey. It cannot go stale between onboardings, because nobody maintains it by hand.
A fix list your agent can work
Ranked by impact ÷ effort, briefed with the rule and the location, served over MCP. Your agent stops burning turns on formatting while a real problem sits there.
Where it fits your rhythm
A survey you run. Never a gate that blocks you.
Calendar-driven audits catch what PR-only scanning can't: new CVEs, bus-factor risk, the rot that accrues when nobody commits.
On a schedule
Per sprint, weekly, monthly or quarterly — your team's calendar, not a commit hook. It runs in the quiet months too, when nobody commits but the world keeps moving.
In the retro
A trend line per lens, slop falling as a visible team win, and a changelog — deltas, findings closed vs opened, features & fixes by area. The sprint retro writes itself.
Watched daily
New CVEs, leaked secrets, score regressions — surfaced the day they appear, between full surveys.
Before the big moments
Scan on demand before a release, a hand-over, or due diligence — with the same rubric the standing cadence uses.
Code composition
How much of your codebase is slop — and how much is brilliant?
Every scan scores the composition of the whole repo: the share that is genuinely brilliant and worth protecting, the fine middle ground, and the slop — duplication, dead scaffolding, unreviewed generated code. The split is on every report card, and the trend shows it falling.
The whole codebase
What the CAI sees that a linter doesn't.
A linter is blind to architecture, ownership and composition — and to what rots between commits. Watchdog scores the system.
Code health
Complexity, duplication, dead code, method bloat read from the compiled bytecode where the runtime exposes it (.NET IL, JVM), and test quality — does the test actually assert anything?
Architecture
Cycles, layer violations, DDD alignment — and a clickable C4 map for bounded-context systems, coupling drawn in red.
Security & compliance
CVEs (SCA across every ecosystem you ship — NuGet, npm, Maven, PyPI, Go modules, Cargo, Composer, RubyGems, Hex and pub), secrets, SAST posture — findings CWE-tagged, a CycloneDX SBOM with every scan.
Maturity & readiness
Tests, observability, ADRs, deployment signals — how ready the system is to run and to be handed over.
Behavioral analysis
Hotspots (churn × complexity), key-person/bus-factor, knowledge freshness, change coupling — mined from git history. Plus rebuild cost (€) and the slop-vs-brilliant split.
One neutral number
The behavioural signals of a dedicated tool are here too — rolled into one neutral, reproducible number and handed to your AI assistant. We never touch your code.
The whole product, not just the main stack
Most products are not one language: a C# API, a TypeScript front end, a Go worker, a Python data job. That normally means four tools, four sets of rules and no shared answer. Watchdog scores all four against the same rubric and rolls them into one number — with the architecture map drawn across the boundaries between them instead of stopping at each one.
One repo, many services, one view
Watchdog decomposes a monorepo into a CAI per deployable service and rolls them up into one product score — so "the platform is fine, the billing service is not" is something you can see, in a language-mixed estate, without running four tools.
Reproducible measurement
Why not just ask an LLM?
An LLM gives you an opinion
Different every run, sees only the slice in its context window, and can't tell you whether the codebase is better or worse than last sprint.
We give you a measurement
Reproducible — same commit, same score. Trended — sprint over sprint, lens by lens. A fix oracle — hand the findings to Claude Code or Cursor over MCP, and the re-scan proves what landed. A board, a stakeholder or a client can re-run it.
Read-only by doctrine
We never touch your code.
Watchdog scans → you (or your AI) fix → we prove. The chain of custody on every change stays yours.
What Watchdog does
Hands you — and your coding agent, over MCP — the finding, the rule, the rationale, the file and line, and the score-impact. It never commits, never pushes, never opens a fix-PR.
What tools that refactor your code do
Edit, commit and open fix-PRs in their own engine. A measurer that rewrites what it grades can't stay neutral — and you lose chain-of-custody on every change.
How it works
Three steps.
Sign in with GitHub
Install the App on the org you want watched. GitHub access ⇒ Watchdog access.
Add a repo
The first scan runs immediately — a baseline CAI and a full report in minutes. The first full report is free.
Pick a cadence
Every sprint, weekly, monthly or quarterly — plus the daily security watch. Trends accrue on their own.
Put a number on your codebase.
Sign in with GitHub · no card · C#, Java, TypeScript, Python, Go and more · the first full report is €0.