Skip to content
Watchdog
Sign inSurvey a repo free

Behind the number

How Watchdog grades your code.

Every Watchdog survey grades your whole product from 0 to 100, whatever mix of languages it is written in. That score is the Code Assurance Index, an open standard. Behind it sit separate grades for each aspect of the code, from how easy it is to change to how well it is protected, and behind each grade sit the individual measurements taken from your code. The rules for all of them are published before anything is scored.

Sign in with GitHub, GitLab, Bitbucket or Azure DevOps · no card · Rust, Java, C#, TypeScript, Go and more

A quick orientation

What makes Watchdog's grading different.

The same code gets the same score

The score is worked out by fixed rules that are published in advance, never by an AI model's opinion. When your score moves, it is because your code, the rules or the known security risks changed, and the survey shows which.

One set of rules for every language

A product written in several languages is measured as one product, by the same rules throughout, so you get one score that means the same whatever the code is written in. When Watchdog adds a new language, the rules stay the same, so your existing scores don't move.

You can check the calculation yourself

Every rule is published before anything is scored, and every survey carries the evidence behind its score. The program that turns that evidence into the number is open source, so you can work the score out again yourself.

Every survey is complete

Every survey measures everything that applies to your code, including the free first report. What you pay for is how much code is surveyed and how often, never how deep a survey goes.

One question per grade

What each grade looks at, and why it matters.

Watchdog grades separate aspects of your code, each one answering a question you would ask about any codebase, such as whether it is easy to change or whether it is safe to run. The Code Assurance Index calls these grades lenses. Each grade is built from individual measurements, which the standard calls dimensions: how tangled a piece of logic is, for example, or whether a library you rely on has a known security hole. Some grades apply to every codebase, and the others only when your product is built the way they are about. When one doesn't apply, its share of the score moves to the grades that do, so a simple service is never marked down for lacking a design it doesn't need.

Always graded

Code health

How easy the code is to read and to change safely. This grade looks at how tangled the logic is, classes that try to do too much, code copied instead of shared, and work left unfinished: stubs that look complete but do nothing, dead code and "fix later" notes. It also catches a long list of small traps that break software quietly, such as errors that are caught and then ignored. Code that is hard to read is slow and risky to change, so every future change costs more than it should.

Always graded

Architecture

Whether the structure holds up as the product grows. This grade looks at how the parts depend on each other, whether those dependencies point the right way and avoid loops, whether each part has one clear job, and whether the boundaries your design intends are kept in the code. It also reads the history for files that keep changing together with no visible link between them, a sign that a boundary sits in the wrong place. When the structure is weak, a change in one place ripples into many others.

Always graded

Maturity

How well the project explains and looks after itself. This grade looks at the documentation, the design decisions your team has written down and whether the code still follows them, whether names are clear and consistent, and whether comments explain why rather than repeat what. It also reads the project's history: files that change often and are hard to read, parts only one person understands, and code nobody has worked on for so long that the knowledge of it has faded. This is what decides whether a new developer, or a new supplier, can take over.

Always graded

Readiness

Whether the product is ready to run in production and to change without breaking. This grade looks at the tests: how much of the code they exercise, whether they actually check results, and whether they pass reliably. It also looks at the automated checks every change must pass, whether problems can be seen once the product is running, how releases are rolled out and rolled back, backups and disaster recovery, and the libraries the product depends on: how up to date they are, whether their licences fit, and whether any passwords or keys have been left in the code.

Always graded

Security & Compliance

How well the product is protected, and how it handles personal data. This grade looks for passwords and keys committed at any point in the history, libraries with publicly known vulnerabilities (CVEs) or published as malicious, security weaknesses in your own code (found by static analysis, SAST), unsafe settings in container and cloud configuration, and platforms that no longer receive security updates. On the compliance side it checks how personal data is protected: encryption, access control, audit trails, how long data is kept, and whether people can have their data exported or erased as the GDPR requires.

Graded when it applies

Domain Modelling

Whether the code protects the rules your business depends on. It applies when your code is built around a model of the business (domain-driven design, DDD). This grade catches business objects that just hold data for other code to change instead of enforcing their own rules; objects that can be created in a state their rules forbid; the same business rule decided in several places; and business logic mixed with database and web code. It checks that each group of objects that must stay consistent (an aggregate) is saved as one unit, one group per change, and refers to the others only by their ID.

Graded when it applies

Event-Driven

Whether the parts of your system can rely on the messages they send each other. It applies when parts of your system communicate by sending messages or events rather than calling each other directly. This grade checks that a change and the message announcing it are saved together (the outbox pattern), so a crash halfway through can't keep one and lose the other. It also checks that each instruction (a command) has a single owner, with news that several parts need to hear sent as an event. It catches code that handles a message by waiting on a call to another service, which makes it depend on that service being available and fast.

Graded when it applies

Event Sourcing

Whether the stored history your system is built on can be trusted. It applies when your system keeps every change as a permanent record of events instead of overwriting the current state. This grade checks that events are defined so they can't be changed once recorded, and that replaying them always rebuilds exactly the same state, with no clock readings, random values or outside calls mixed in. It flags personal data written into records that can never be deleted, unless it is encrypted in a way that lets it be made unreadable on request (crypto-shredding), because the GDPR gives people the right to have their data erased.

Graded when it applies

Accessibility

Whether your interface is built so that everyone can use it, including people who can't see the screen or can't use a mouse. It applies when your code includes a web interface. This grade checks for text alternatives on images and captions on video, labels on forms and buttons, a clear page structure, keyboard access to everything, enough colour contrast, respect for people who switch off animation, correct use of the attributes screen readers rely on (ARIA), and automatic accessibility tests in your build. Because it reads the code, it covers the parts of the accessibility guidelines (WCAG) that can be checked there, and shows how ready the interface is rather than certifying it.

Graded when it applies

Performance

Whether your code is written and measured with speed in mind. It applies when your product's design calls for it. Benchmarks that measure the code's speed as it changes, and code written to use memory sparingly, for example by working on data in place instead of copying it, earn credit where they are present and cost nothing where they are absent. The grade marks down code that sits blocked while it waits for other work to finish, which wastes capacity and can make the program hang.

Every individual measurement is listed in the standard's catalogue, with what it checks and which grade it belongs to. How a measurement is taken, and how certain its result is, is explained in the documentation under How we measure.

From lenses to one number

How the lenses become one score.

Each lens score is worked out from the measurements taken inside that lens, and the lens scores are then combined into one score for your whole product. Every step is set by the published rules, so anyone holding the survey can trace the calculation from each measurement to the final number.

Some lenses count for more

The lenses don't count equally. The rules set how much each lens counts and publish it with every version, so you can see what each lens contributed to your score. That is also why fixing a problem in one lens can move your score more than fixing a similar problem in another.

The same rules for every language

A product written in several languages is surveyed as one product, not scored one language at a time and averaged. Every file is measured under the same rules and counted in the same lenses, so a score of 74 means the same thing whatever the code is written in. When Watchdog adds a new language, the rules stay the same, so the scores you already have don't move.

When a lens can't be measured, the survey says so

Some lenses can be measured more fully in some languages than in others. Where a lens can't be measured properly, the survey marks it as not measured and gives the reason, and its share of the score moves to the lenses that could be measured, just as it does for a lens that doesn't apply. The language you chose never lowers your score, and a gap in what could be measured never counts as a zero.

Your code's history counts too

Some measurements come from the history of your code rather than from the files as they are today: files that change often and are hard to read (hotspots), parts that only one person understands (bus factor), code nobody has worked on for so long that the knowledge of it has faded (knowledge freshness), and files that keep changing together with no visible link between them (change coupling). They are measured by the same fixed rules and count towards the same score as everything else, because they show things that reading today's files can't.

One failing lens limits the whole score

A lens in the lowest band, Critical, limits the whole score, however well the other lenses do. The stronger lenses still count and still show in the number, but they can't make up for the one that is failing, so a serious problem in one place always shows in the score.

What a minimum score means

A minimum score in an agreement can be checked lens by lens. A minimum of 80 means every lens that always counts is Strong or better and no lens is Critical, and that meaning is fixed when the contract is written, so both sides know what the number stands for.

How the standard calculates the score →

Same code, same score

Only three things can move your score.

A score you follow over time, or write into a contract, is only useful if you know why it moves. Every Watchdog score is calculated under a published set of scoring rules (the rubric), and every survey records the version it used, so when your score moves, the survey shows which of these three things caused it.

A change in your code

The score comes from tools that follow fixed rules and give the same answer every time they read the same code. That is what makes a change in the score mean something: when you fix a problem the score can rise, and when new code brings new problems it can fall. Each survey lists the findings behind its score, so you can see where a change came from and whether the work your team put in moved the number.

A new version of the rules

Any change to the rules that could move the score of unchanged code becomes a new, named version, published with a list of what is different, so when you compare your scores over time, you can always see whether the rules changed in between. Every survey names the version it was scored under, and a contract can lock one version for its whole term: the score agreed at the start and the score at the end then measure the same thing, and neither side can be surprised by a rule that arrived in between. Our own work is held to this too. A change to the scoring can't be released unless the published rules say the same thing, and a test codebase that never changes is scored again and again, so any update that would move its score without a new version is stopped before it is released.

How the rules are versioned →

A newly published security hole

Security risks change even when your code doesn't. When a security hole in a library you use is made public, a security finding can change even under a locked version of the rules, because the risk to your product is real from the day it becomes known. The survey records it separately from any change in your code or in the rules.

Watchdog never lets an AI model decide your score. An AI model can give a different answer each time it is asked the same question, so a score that relied on one could change even though your code had not. That is the difference between a measurement and an opinion, and a Watchdog score has to mean the same thing every time you read it, whether you are tracking your own progress or holding a supplier to a contract. Where Watchdog uses an AI model, it explains findings in plain words beside the number. Even those explanations are checked: one that falls outside set limits is not used, and one that lands on the borderline is run again.

Your own inputs can't move it either. Compliance declarations, findings you choose to suppress and the terms of a contract change what the report says, never the score. Two parties can agree between themselves how a finding should be presented, and the number they are both looking at stays where the evidence put it.

Every rule a score was worked out under is published. Every survey carries the evidence behind its score, and the program that turns that evidence into the number is open source. Anyone holding a survey can run it and get the same number back. That confirms the calculation; how each measurement was taken is set out in the survey itself.

Check the calculation yourself →

How to read the number

Every score is read on the same fixed scale.

Every score is shown on the same scale, from worst to best: Critical, Weak, Adequate, Strong and Exemplary. The number itself is the reading, and the band is a name for where it falls, so a score can be talked about in a sentence. The scale is the same for everyone and is not a ranking against other codebases: the lines between the bands are set in the published rules, and they are shown below exactly as the rules state them.

Each band that has a published survey shows a real one you can open and check. Where a band has none, it is because a repository's published survey is its best so far: a survey that scores below one already published is never published, so the bottom of the scale holds few published surveys, or none.

Exemplaryfrom 90

ruben-rasmussen/auth98

Adequatefrom 50

aframevr/aframe70

Criticalunder 25

No published survey in this band.

Findings you can act on

Every measurement is tuned to find real problems without flagging everyday code.

A finding you can't trust is worse than no finding: someone has to spend time checking it, and every wrong one makes the rest of the list harder to believe. So each measurement is tested on real code, both on the problems it is meant to catch and on the everyday patterns that only look like one. For example, code that nothing in your product uses is normally reported as unused (dead code), but not when it is one of a fixed set of functions your code has to provide (an interface). All of this is done by Watchdog, so there are no rules for you to set up or adjust.

Every language must pass the same test

A language is added only after a set of real projects written in it has been surveyed and the findings have held up when checked. Its measurements are tuned on code in that language alone, never on another language's, even when two languages look alike. The surveys published on this site are open to read, so you can judge the findings for yourself on a project you already know.

What happens when you dispute a finding

You can dispute any finding that counts towards your score and give your reason, and a person at Watchdog reviews it and decides. A dispute never changes your score or removes the finding. If the finding was wrong, the measurement itself is corrected, and the case is added to its tests to stop the same mistake from happening again. If the dispute is declined, the finding keeps a note explaining why it stands, and that note is the answer to anyone who raises the same point again.

The fixes that matter most come first

Each survey comes with a list of what to fix, ranked by how much each item is expected to raise your score for the work it takes. The score weighs each kind of problem by how widespread it is, so a single stub that looks finished but does nothing barely changes it, while the same problem repeated across the code does.

You can tell a clean result from an unchecked one

A long list of findings doesn't prove a survey was thorough, and a short one doesn't prove the code is clean, because it may only mean that fewer measurements could run. That is why the report states how many of the measurements that apply to your code returned a result, and marks the rest not measured, with the reason.

Worked example: a class that only looks like it does too much

In many languages, code is organised into classes: units that each hold some data together with the functions that work on it. When a class's functions fall into separate groups, and no group uses the data or functions of another, the class is usually doing several jobs at once, which makes it hard to change safely. A measurement called LCOM4 counts those groups, and Watchdog uses it in the Architecture lens to find classes like that.

LCOM4 gets some good designs wrong. Take a class for an order in an online shop, with functions to add an item, change the shipping details and cancel the order (Order, with AddItem, ChangeShipping and Cancel). Each function works on different data, so LCOM4 sees three separate jobs. Yet the three belong together: between them they keep the order's business rules true (its invariants), so the order can never end up in a state the business doesn't allow. To the count, this well-designed class looks like one that does too much (a god class).

Watchdog recognises the designs LCOM4 gets wrong and leaves them out of the count: business objects like the order that protect their own rules (aggregates), classes whose functions all store and fetch one kind of record (repositories), code that a tool writes for a screen (generated view models), a shared class that holds what several other classes have in common (a base class), and functions a class must have because an interface requires them. A class that really does too much is still reported.

Because Watchdog recognises these designs by how the code is put together, the recognition works the same way in every language where this measurement runs, and the Languages page shows the languages where it doesn't.

The tuning is never finished. Code that relies heavily on the conventions of its language or framework can still produce some findings that are wrong or unhelpful, and each round of tuning reduces them.

How Watchdog checks its own findings →

Code that only looks finished

What the survey finds in code that compiles and passes its tests.

Code written fast and in large amounts is still checked by the compiler and by style checkers (linters), which catch mistakes in how it is written. They don't catch code that does nothing, hides a failure, or makes every later change harder, and the tests can pass while all of it is there. The survey looks for these problems in the code itself, so the same problem gets the same finding whoever wrote it.

Code health

Unfinished functions

A function called CalculateTax that takes an order and always returns 0 compiles and can pass its tests, yet no order is ever charged tax. The survey recognises these placeholders (stubs) by how the code is built rather than by what it is called: functions that only report that they haven't been written yet, functions that take inputs but always return the same value, functions marked as waiting for slower work that never wait for anything, branches that can never run, and classes whose functions are mostly placeholders.

Readiness

Tests that check nothing

A test that runs the checkout code but never checks the total passes whether the total is right or wrong, and the measure of how much code the tests run (coverage) still counts that code as tested. The survey looks for tests that check no result, and for tests that are switched off but still make the test suite look bigger than it is.

Code health

Errors that disappear

When a payment fails inside code that catches the error and then does nothing with it, the program carries on as if nothing went wrong: no error message, no log entry, no second attempt. The survey finds errors that are caught and ignored, and errors passed on in a way that loses the record of where they started (the stack trace), which makes the cause harder to find.

Code health

Notes to fix later, and code nothing uses

A comment that says "TODO: check the input" marks work someone knew was missing. The survey counts these notes (TODO, FIXME and HACK), warnings that have been switched off in bulk, code that has been commented out, and code that nothing uses. Each one is small, but together they count towards the score, so you can see how much known work is waiting and whether it is getting smaller.

Code health

Code copied instead of shared

When the same block of code is pasted into several places, a bug fixed in one copy stays in the others, and every change has to be made more than once. Watchdog finds blocks that are identical or nearly identical and scores them by how much of your code they make up. Code generated by tools is left out, because that repetition comes from the tool, not from your team.

The checks for unfinished functions and for errors that disappear read C# code so far, and the check for tests reads several languages. Where a check can't read your code, the report says so and that check isn't scored.

Our own findings, checked in public

How Watchdog measures the noise in its own findings.

A finding counts as noise if it is wrong, states an opinion as if it were a fact, repeats another finding, or doesn't apply to the kind of software being checked. Those are the findings that someone who knows the code well would not act on. Watchdog's target is that fewer than 5% of the findings that count towards your score are noise. No measured result exists yet, and none will be quoted here until a measurement has passed every check described below.

The method is published first

Everything about how the rate is worked out is published before any measurement starts, so none of it can be adjusted to fit the result: what counts as noise, how the rate is calculated, how the projects are chosen and how findings are judged. The version of Watchdog under test stays fixed from start to finish, so the figure describes one exact version.

Nobody chooses the projects

The projects for each round of measurement are picked by a random draw whose starting number (the seed) is generated rather than chosen, and announced before any results exist. The only condition for being in the draw is that a project can be surveyed without errors. Anyone can repeat the draw and get the same projects, which confirms that Watchdog didn't pick ones it does well on.

People check the judging

An AI model judges every finding against the code itself, following instructions anyone can read word for word. A sample of the findings is also judged by two people, separately and without seeing the model's answers, and at least one of them did not build Watchdog. If the model disagrees with them too often, the instructions are revised and every finding is judged again, and when a disagreement can't be settled, the finding counts as noise.

Everything behind a result is published

Before each round, Watchdog commits in public to publishing the result, whether it meets the target or misses it. The result comes with every finding, every judgement from the model and from the people, how often they agreed and the statistics, along with a small program anyone can use to recalculate the rate. The number of findings per thousand lines of code is published beside the rate, because a tool that reports almost nothing would look accurate without being useful.

Read the published method

Live from a real survey

Your architecture, drawn from the code.

The Architecture lens looks at how your whole system is put together, rather than one file at a time. When your code is organised into separate areas (bounded contexts), the survey draws a map of them from the code itself: the main parts of your product and the dependencies between them, in a common format for drawing software architecture (the C4 model). Each part is coloured by the number of findings per file, and a red line marks a place where a business type that belongs inside one part is used by another, which ties the two together. Because the map is drawn again from the code on every survey, it always matches the code the survey read. Below is the map from the published survey of Kubernetes, drawn the same way the app shows it.

Health by finding density · red = boundary coupling (D23)k8sio/api[namespace]health 100 · 0 findings · 120…k8sio/apiextensions-apiser…[namespace]health 100 · 0 findings · 76 …k8sio/apimachinery[namespace]health 100 · 0 findings · 214…k8sio/apiserver[namespace]health 100 · 1 finding · 555 …k8sio/cli-runtime[namespace]health 100 · 0 findings · 42 …k8sio/client-go[namespace]health 100 · 0 findings · 227…k8sio/code-generator[namespace]health 100 · 0 findings · 279…k8sio/component-base[namespace]health 100 · 0 findings · 98 …k8sio/component-helpers[namespace]health 100 · 0 findings · 40 …k8sio/dynamic-resource-all…[namespace]health 100 · 0 findings · 52 …k8sio/endpointslice[namespace]health 100 · 0 findings · 11 …k8sio/kubectl[namespace]health 100 · 0 findings · 202…k8sio/kubernetes[namespace]health 100 · 1 finding · 2060…k8sio/pod-security-admissi…[namespace]health 100 · 0 findings · 28 …
How the map is drawn, and how to read it →

See how your own code is graded.

Sign in with GitHub, GitLab, Bitbucket or Azure DevOps · no card · Python, Kotlin, Java, Go, C# and more · the first full report is free.