Skip to content

tepyd cover

Which tier actually executes each unit, and is it the cheap one?

cover is the only lens that runs your test suite. For each tier it does one coverage run -m pytest <tier-dir>, dumps coverage json, and measures, per source unit, what fraction of its statements that tier actually executed. Because the tiers partition the suite, the total cost is roughly one full run.

$ tepyd cover                  # all tiers
$ tepyd cover --tier a_unit    # just one tier (repeatable)
$ tepyd cover --json
$ tepyd cover --min-src 5      # by source *statements*, default 20

Running it from the right environment

cover imports and executes your code

Unlike the static lenses, cover runs your suite under the same interpreter that runs Tepyd. That interpreter must therefore have your project's dependencies, plus pytest and coverage, importable.

Add Tepyd to the project you are analysing and run it from there:

$ uv add --dev tepyd
$ uv run tepyd cover

Installing Tepyd standalone (uv tool install) and pointing it at the project with -C will fail to import your tests. If it can't import your project, cover says so, names the interpreter and the missing module, and points at the fix. Earlier versions printed a wall of zeros.

Output

unit      stmts  unit  e2e   any
--------  -----  ----  ---  ----
checkout      7    0%  88%   88%  v hidden
domain        9  100%   0%  100%
stmts
The unit's source statements, as coverage.py counts them. (Statements, not LOC, so --min-src here is a different threshold from shape's.)
one column per tier
The fraction of those statements that tier's tests executed.
any
The union across tiers: the unit's true reachable-by-tests coverage.

The hidden inverted pyramid

The v hidden flag is the headline diagnostic: a unit that is well covered overall (any high) but barely by the unit tier. The lines run, but only the expensive tiers run them.

A global coverage report would show both rows above as green. Only this lens reveals that checkout's coverage is entirely end-to-end: slow to run, imprecise when it fails, and brittle against unrelated changes.

Measured zero, or not measured

cover goes out of its way to distinguish "measured zero" from "did not measure":

  • It prints per-tier progress to stderr as it runs. It executes the whole suite once per tier, so a large suite takes a while. The progress lines show it has not hung.
  • It ignores the project's own [tool.coverage] config, so the numbers do not depend on it, and attributes coverage by resolved path, which holds up for multi-file units and absolute coverage paths.
  • A tier that fails to run is shown as a 0 % column and listed as not measured, distinct from "0 % because untested".
  • A tier whose tests ran but failed is used, with a warning that the numbers are a floor.
  • A tier that runs but measures 0 % everywhere, typically a browser/Playwright suite driving your app in a separate process coverage.py cannot see, is flagged as not measured rather than "covers nothing". Scope it out with --tier, or measure it under subprocess coverage.

Failures and warnings print to stderr even on the success path, so they are not lost when stdout is piped to a file. The exit status stays 0; --json includes the full detail.

--json

{
  "tiers": [{ "name": "a_unit", "label": "unit" }],
  "units": [
    {
      "unit": "checkout",
      "statements": 7,
      "tiers": { "a_unit": { "covered": 0, "pct": 0.0 } },
      "any": { "covered": 6, "pct": 0.8571 },
      "hidden_inverted": true
    }
  ],
  "failures": [{ "tier": "c_e2e", "reason": "..." }],
  "warnings": [],
  "blind_tiers": ["e2e_playwright"]
}

failures and blind_tiers are what a CI consumer needs to tell a broken measurement from a real zero.