Run it yourself
A scanner built for the languages you actually use
The hosted survey reads your repository on our hardware. The CLI does a narrower job on yours: you pick the languages your codebase is written in and the platform you run on, and we compose a binary that contains exactly those analysers and nothing else.
That is the whole reason it is small. A build carries the language models you chose, for the one platform you chose — not every language we support, and not the native libraries for ten operating systems you do not use. Each build is cached, so the second person to ask for the same combination gets it immediately.
Build yours
Choose one or more languages and a platform. Anything you leave out is still detected — the scanner reports it as code it could not read, rather than staying quiet and letting a clean result stand for a file nothing looked at.
- 1 Which languages does your codebase use?
Anything you leave out is still detected and reported as unread. A scanner that stayed quiet about code it cannot parse would be reporting a clean result about code nothing looked at.
Each language shows how many of the dimensions in your build can read that language. Each dimension declares that for itself and the table is generated from those declarations, so this figure cannot be typed into the page. It is what the scanner is able to analyse, not a promise about your repository: a dimension that reads Dockerfiles still needs you to have one.
C# 13 of 13 read it Dart 13 of 13 read it Elixir 13 of 13 read it Erlang 13 of 13 read it F# 13 of 13 read it Go 13 of 13 read it Java 13 of 13 read it Kotlin 13 of 13 read it PHP 13 of 13 read it Python 13 of 13 read it Ruby 13 of 13 read it Rust 13 of 13 read it Scala 13 of 13 read it Swift 12 of 13 read it TypeScript / JavaScript 13 of 13 read it VB.NET 13 of 13 read it - 2 Which platform will it run on?
One binary, one platform — that is most of why it is small. A portable build carries the native libraries for every other platform too.
Linux (x64) Linux (arm64) macOS (Apple silicon) macOS (Intel) Windows (x64) Windows (arm64) - 3 Do you already have these scanners?
Six dimensions call tools we do not ship. Tick only the ones already on the machine that will run this — anything you leave unticked is left out of the build entirely, so everything your scanner carries is something it can measure. The README in your zip names the install commands for whatever you did tick.
gitleaks adds D28 osv-scanner adds D30, D43 semgrep adds D29, D32 trivy adds D31
Pick at least one language to see your scanner.
Your scanner Build this scannerBuilt for your selection and cached, so the next person asking for the same one gets it straight away. The size shown is this build's own — measured from the archive, not an estimate. The full hosted engine is 131 MB.
✓ What it measures — 13 of 165 dimensions
This is not the full survey. The hosted engine runs 165 dimensions; the CLI carries the 19 that are decidable from a checkout alone — no git history, no runtime evidence, no judged pass. The rest are withheld by construction: the binary does not contain them, so it cannot quietly skip them either. Every run writes a notice naming exactly what it shipped and what it left out. The 6 marked needs a tool are only built in if you ticked step 3.
- D1 Cyclomatic complexity
- D2 Cognitive complexity
- D12 Dependency hygiene
- D13 Secret scanning
- D14 Licence compliance
- D28 Secrets in history needs gitleaks
- D29 Static analysis needs semgrep
- D30 Dependency vulnerabilities needs osv-scanner
- D31 Infrastructure security needs trivy
- D32 Data compliance needs semgrep
- D43 Malicious dependencies needs osv-scanner
- D44 Platform end-of-life
- AC1 Text alternatives
- AC2 Forms and labels
- AC3 Page structure
- AC4 Keyboard semantics
- AC5 ARIA correctness
- AC6 Visual and motion
- AC7 Accessibility enforcement
What it measures, and what it does not
The CLI carries a small, fixed subset of the engine's dimensions, and a default build carries fewer still. It is not the full survey, and it should not be read as one.
The subset is the dimensions decidable from a checkout alone — no git history, no runtime evidence, no judged pass. Some of those shell out to scanners we do not bundle, so they are left out unless you tell the builder above which ones you already have. The rest of the engine needs things a local binary does not have: a repository's history, a running deployment, a second opinion, or a corpus to compare against. They are not disabled in your build; they are not in it. The binary cannot quietly skip a dimension it does not contain, and every run writes a notice naming exactly what it shipped and what it left out.
The builder above states the exact counts, for the languages and scanners you pick — they are read from the engine rather than written here, so they move when it does. A green CLI run is evidence about the questions your build actually carried; the notice names them. It is not a Code Assurance Index score, and we do not publish one from it. If you need the number, you need the hosted survey.
Languages you did not choose
Detection is separate from analysis, which is what makes the disclosure possible: the scanner can see a language it cannot read. If your repository turns out to be mixed, the run says so — how many lines, what share of the tree — and recommends the combination that would cover it. A team that adds a service in a new language six months from now finds out from the scan, not from remembering this page.
Using it from a terminal
Unzip it and point it at a checkout. The .NET runtime is bundled, so you do not need .NET installed, and there is nothing to log into — the binary reads the working tree and writes its output beside it.
Some dimensions shell out to scanners we do not bundle — semgrep, gitleaks, osv-scanner and trivy. Without them those dimensions abstain, and the run says so rather than reporting a clean result. The README.md in your zip names exactly which ones your build needs, because that depends on what you selected.
unzip watchdog-cli-go-linux-x64.zip -d watchdog-cli
chmod +x watchdog-cli/codehealth
# the path comes FIRST — there is no sub-command
watchdog-cli/codehealth . --output ./watchdog-out
# what this binary contains, and what it withheld
cat ./watchdog-out/edition-notice.mdIt exits non-zero when a finding trips the bar you set, so it is usable as a gate with no further wiring. Read edition-notice.md at least once: it is the file that tells you which dimensions this particular build carries, and which languages it found but could not read.
If a dimension ran but could not measure, the run says so — on the console, in the notice, and in the pull-request comment. A scanner whose toolchain fails emits no findings, which looks exactly like a scanner that looked and found nothing. That distinction is not left to you to notice.
Using it with a coding agent
The SARIF output is the point of contact. An agent that can read a file can read the findings, and each one carries the file, the line and the rule that produced it — so the agent is fixing something located, not guessing from a summary.
# produce findings the agent can read
watchdog-cli/codehealth . --output ./watchdog-out
# hand it all three, not just the first
# report.sarif what was found, with file and line
# edition-notice.md what this build did NOT look at
# pr-comment.md the same run, written for a human readerGive it the notice as well as the findings. An agent handed only a findings file will reasonably infer that everything else passed. The notice is what stops it concluding that your Kotlin service is clean when the binary never had a Kotlin analyser in it.
The same applies to every dimension the CLI does not carry, and to any your build left out. If the agent is going to act on "no findings", it needs to know the question was narrower than the whole survey.
Using it in a pipeline
Cache the binary rather than fetching it every run — it is built for one commit of the engine and does not change between promotes. Most CI systems will also ingest the SARIF directly and annotate the diff with it.
# GitHub Actions
- name: Watchdog CLI
run: |
curl -sSL "$WATCHDOG_CLI_URL" -o cli.zip
unzip -q cli.zip -d cli && chmod +x cli/codehealth
cli/codehealth . --output watchdog-out --fail-on issue
- name: Comment on the pull request
if: always() && github.event_name == 'pull_request'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: gh pr comment ${{ github.event.number }} --body-file watchdog-out/pr-comment.md
- name: Publish findings
if: always()
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: watchdog-out/report.sarif
- name: Keep the disclosure with the run
if: always()
uses: actions/upload-artifact@v4
with:
name: watchdog-edition-notice
path: watchdog-out/edition-notice.mdPublish the notice as an artifact too. A pipeline that stores only the findings loses the record of what the run could not see — and six months later nobody can tell whether a quiet period meant the code was clean or the scanner had no analyser for it.
Commenting on the pull request
Every run writes pr-comment.md next to the SARIF: the verdict, the worst few findings with their file and line, and what the build could not see. Your pipeline posts it with its own credential — one line, as in the step above.
We render it; you post it. That split is deliberate. For us to post the comment ourselves, this binary would have to carry a token with write access to your repository, on every runner you copy it to. Handing you a file instead keeps that token where it already is, works the same on GitHub, GitLab and Azure DevOps, and leaves we never write to your repository true.
The comment states which dimensions the build carried and which languages it could not read even when the run is clean — which is exactly when "no findings" is most likely to be read as "nothing wrong" rather than "nothing this build looked at". A comment that discloses on a noisy run and goes quiet on a green one teaches a team that silence means safety.
Gating on it
A non-zero exit is the gate. Start by failing on the dimensions you already act on rather than everything your build carries at once — a gate that fires on everything from day one gets switched off in a week, and a gate nobody trusts is worse than no gate.
Questions worth asking first
Is this the same engine as the hosted survey?
Yes — the same analysers, composed differently. A dimension the CLI carries gives the answer the hosted survey would give for that dimension. What differs is how many dimensions there are, not what each one says.
Does it send anything to you?
Nothing goes to us. It reads the working tree and writes to the output directory you name. There is no account, no telemetry and no upload — it is a binary on your machine, not a client for our service.
It does reach public package registries, and you should know that before running it on an air-gapped or audited machine. Licence and vulnerability dimensions look your dependencies up by name — crates.io, npm, NuGet, PyPI and so on. Only the package name and version leave the machine; your source never does. On a large dependency tree this is also the slowest part of a scan: one Rust repository took eight minutes, of which five were spent pacing 424 crates.io lookups.
How long does the first build take?
Under a minute for most combinations, because it is a real compile. It is cached afterwards, so the same selection is immediate for everyone who asks next — and a new build is composed whenever we promote a new engine, so the binary you get is never older than the analysers behind it.
Can I get a build for a combination that is not offered?
The languages and platforms in the picker are what the generator knows how to compose. A language is selectable only where we can actually ship the analyser frontend it needs — the rest are shown, greyed, and marked not available yet, because a scanner that arrives unable to read the language you picked would report nothing and look clean. If the one you need is greyed out, ask us; the list is a fact about what is built, not a commercial boundary.
Can it replace the hosted survey?
For the questions your build carries, on a developer's machine, yes. For the Code Assurance Index, a trend, evidence a counterparty can verify, or anything needing history or a running deployment, no — those are the dimensions that are not in the binary.