Skip to content
Watchdog
Sign inSurvey a repo free

For engineering teams

Know the state of your code, every sprint.

An engineering team knows what it shipped this sprint. What it rarely has is an overview of the whole codebase. There is no starting point to measure progress against, and no clear view of where the code has become hard to change, which makes every new change slower (technical debt). Without that overview, it is hard to show where time spent improving the code would help most, or whether the last cleanup made a difference.

Watchdog surveys the whole product on the team's schedule, such as every sprint, without adding a step that every change must pass. It grades the code from 0 to 100, whatever stack it is built in. The first survey gives the team its starting point, and every survey after it shows what has changed, including in the architecture and security. Each survey also ranks what to fix first by how much each fix lifts the score for the work it takes. The team's developers, or their coding agents, work through the list, and the next survey shows whether the fixes worked.

Sign in with GitHub, GitLab, Bitbucket or Azure DevOps · no card · nothing to install · read-only access · the first survey of every repository you connect is free.

Every sprint

a survey of the whole product on the team's schedule, with nothing added to the checks every change must pass

Daily

a security watch on the libraries you depend on, passwords and keys left in the code, and unsafe settings, so the team hears about a new problem the same day it is found

Same code

always gets the same score, so the team can compare one sprint with the next

€0

for the first survey of every repository you connect, with nothing held back

From the developer to the lead

What the team knows after every survey.

The same survey serves the people who write the code and the people who plan the work. Developers, testers and the people who keep the product running see where the code is hardest to follow or most at risk, with the most serious problems first. Tech leads and engineering managers see where to point the team this sprint, and whether last sprint's work raised the score. Tools that check each change as it is made are good at catching problems in the lines that changed. Watchdog works alongside those tools and looks at the whole product: how its parts fit together, who knows which parts of it, and how it has changed since the last survey.

An answer to "is it getting better?"

Every survey grades the code by the same published rules, so the scores from one sprint and the next can be compared. The team sees the score for the whole product and for each area of it, such as security or architecture, and whether each one moved up or down. Each score can be opened to see the findings behind it and the rule each one is measured by, so the team can discuss in the retro exactly what changed in the code and why the score moved.

A case for the refactor

Time to improve the code is hard to plan when the only argument is that the code feels hard to work in. A survey shows which part of the code has got worse over the last several surveys, with the file and line of every finding behind it. That gives the team a concrete reason to set time aside for the work, next to new features. When the work is done, the next survey shows whether the score went up, so the team can show that the change improved the code, not just that it still compiles. That result makes the next refactor easier to agree on.

An architecture diagram that matches the code

When a product is divided into separate areas, such as orders, payments and users, every survey draws a diagram of them from the code itself, showing how they depend on each other and where the boundaries between them are. Areas that are tied together where they shouldn't be are marked in red, and each area is coloured by how many problems the survey found in it. Nobody has to keep the diagram up to date by hand, so it is as current as the last survey, including for new team members. The architecture is graded in every survey, with or without a diagram.

A fix list for the team and its coding agents

Every survey ranks its findings by how much each fix is expected to raise the score for the work it takes, so the first items on the list give the most improvement for the least work. The survey also names the weakest area and what would raise it, which gives the lead an order for the backlog that is easy to explain. Each finding comes with the rule it breaks, the file and the line. Developers can work through the list themselves, or give it to a coding agent through the standard way agents connect to tools (MCP). The agent then spends its time on the problems that matter most, instead of on small things like formatting.

One score for the whole product

Most products are built in more than one language, such as an API in C#, a front end in TypeScript, a background worker in Go and a data job in Python. Checking them usually takes a separate tool for each language, each with its own rules, and the results are hard to add up to one answer for the product. Watchdog grades every part by the same rules and combines them into one score. Repositories that are shipped together can be surveyed as one product, and the architecture diagram shows how the parts connect across languages. When several services live in one repository, each service also gets its own score, so the team can see that the product as a whole is fine while the billing service is not.

Where the team depends on one person

The survey reads the history of the code and finds the parts that only one or two people have worked on recently (the bus factor). Those parts would be the hardest to take over if someone left, and knowing where they are lets the lead plan for it, for example by having a second person work on them. The same history shows the files that change often and are hard to read (hotspots). Those are the files the team works in most, so problems there cost the most. The survey also gives a rough estimate of what it would cost to build the code again from scratch, in euros and working years, which helps the lead explain what it would take to replace it. All of this comes with the survey, so no separate tool is needed to read the history.

Every area the survey grades, and the rules behind each score, are explained on What we measure. Every survey also includes a list of every library and licence in the code (CycloneDX SBOM), and each security finding carries its standard weakness code (CWE).

Live from real surveys

How much of the code is worth protecting, and how much needs work?

Every survey scores the code file by file, from 0 to 10, by what it found in each file, and divides the files into three groups. Files marked brilliant are well organised, tested and worth protecting. Files marked fine are the ordinary middle. Files marked slop contain copied code, leftover code that nothing uses, or unfinished placeholders. The split is in every report, and each new survey adds a column, so the team can see whether the share of good code is growing. The three projects below are real open-source projects surveyed by Watchdog.

lau/tzdata

  • 18 Sep 2026: 28.6 % brilliant, 57.1 % fine, 14.3 % slop
  • 3 Oct 2026: 28.6 % brilliant, 42.8 % fine, 28.6 % slop

oldest2 scanslatest

28.6 % brilliant42.8 % fine28.6 % slop

latest scan · 7 files scored

georgeguimaraes/arcana

  • 18 Sep 2026: 37.5 % brilliant, 58.3 % fine, 4.2 % slop
  • 3 Oct 2026: 36 % brilliant, 60 % fine, 4 % slop

oldest2 scanslatest

36 % brilliant60 % fine4 % slop

latest scan · 25 files scored

marciok/gust

  • 18 Sep 2026: 40 % brilliant, 40 % fine, 20 % slop
  • 3 Oct 2026: 50 % brilliant, 30 % fine, 20 % slop

oldest2 scanslatest

50 % brilliant30 % fine20 % slop

latest scan · 10 files scored

Alongside the team's AI tools

Why the team can trust the score.

Many teams already ask an AI model what it thinks of their code, and use coding agents to change it. Watchdog has a different role in that work: it measures the code by published rules, and it leaves every change to the team.

A measurement, not an AI model's opinion

An AI model's answer can differ each time it is asked. It also reads only as much of the code as fits in one request, so it cannot say whether the whole product got better or worse since the last sprint. Watchdog grades the whole product by fixed, published rules. The same code gets the same score every time, and the score moves only when the code, the rules or the known security holes change. An AI model never decides the score, and where one is used, it only helps explain a finding. A board, a stakeholder or a client can work the score out again from the evidence. When code arrives faster than the team can review it, a separate page explains what the survey checks.

Watchdog never changes the code

Watchdog gives the team, and coding agents such as Claude Code or Cursor, each finding with the rule it breaks, the reason, the file and line, and how much fixing it would raise the score. It never commits or pushes code, and never opens a pull request. The team or its agent makes the fix, the next survey checks it, and every change to the code stays the team's own. Because Watchdog never changes the code it grades, it has no stake in the result, and the score stays neutral.

How coding agents use the findings →

How it fits in

How the survey becomes part of every sprint.

Checks that run on each change only run when someone changes the code. Some risks arrive without any change: a security hole is found in a library the product uses, or the knowledge of a part fades because nobody has worked on it for a long time. A survey on the team's schedule, and a security watch every day, catch those as well.

Connect your code host

Sign in with GitHub, GitLab, Bitbucket or Azure DevOps. Watchdog only reads the code, and it sees only the repositories the team gives it access to.

Get the starting point

Add a repository, and the first survey starts at once. It takes minutes for a small project and longer for a large one. The first survey of every repository is free, and its report is the starting point that every later survey is compared with.

Survey after every sprint

Place the repository on the team, and a survey runs after each of the team's sprints, whether they last a week, two weeks or a month. It also runs in quiet periods when nobody changes the code, so the risks described above are still caught. Between surveys, a daily security watch checks the libraries the product uses, passwords and keys left in the code, and unsafe settings, and sends what it finds to Slack, Microsoft Teams or Discord.

Use each report in the retro

Each report shows how the score of the whole product and of each area moved, how many findings were closed and how many opened, and how the share of good code changed. It also includes release notes drafted from the code's history, and a sprint plan that ranks what to fix first. Before a release, a handover or due diligence, the team can run an extra survey by the same rules.

Put a number on your codebase.

Sign in with GitHub, GitLab, Bitbucket or Azure DevOps · no card · Java, TypeScript, C#, Python, Go and more · one fixed price for the whole team, however many developers it has · the first full report on every repository is €0.