1. The metrics I actually care about
Lead with these. The rest are supporting data, not the headline.- Deployment frequency — how often we ship to users. Drives every other downstream signal.
- Lead time for changes — commit to production. Shortening this is the single biggest health indicator for me.
- Change failure rate — what % of deploys cause a rollback or hotfix. The trust number.
- PR merge time — from open to merge. The review-bottleneck detector.
2. The metrics I treat as supporting context
Useful, but only when the headline metrics move:- PR throughput — sanity check that velocity is consistent
- Commit frequency — early signal for upcoming throughput change
- Hotfix rate, revert rate — diagnostic when change failure rate spikes
- Average PR size — diagnostic when merge time creeps up
3. My healthy-vs-concerning thresholds
Numbers above are defaults. Override via memory if a team has a different baseline I’ve explicitly set.
4. What I want flagged
Surface only what moved or what’s out of healthy range. I don’t want a wall of numbers.- Significant week-over-week swing — anything moving ±20% or more
- Anything crossing into “Flag immediately” thresholds
- Wins worth naming — e.g. deploy frequency doubled, lead time halved
- Suspicious zeroes — zero deploys could mean the release mechanism changed, not that we stopped shipping. Say so.
5. How I want the report shaped
If CORE produces a report:- One flat table, one row per period. Easy week-over-week scan.
- Key Observations section underneath — capped at 3–5 bullets. Lead with the metric name, then the insight. Don’t list every metric.
- No padding. If nothing notable changed, say “all metrics stable” and stop.
The same table shape feeds Google Sheets (when connected), so each run appends a row and the spreadsheet becomes a historical chart over time.
6. What I don’t want
- Don’t list every metric every time. Lead with what moved.
- Don’t editorialise without a number — every observation must point at a specific metric and direction.
- Don’t manufacture concern from a single data point. One bad week is noise; a 3-week trend is signal.
How this lens is used
Other surfaces consume this lens:- Weekly engineering report task applies the lens to last 7 days, logs to Sheets
- On-demand query (“how was last week?”) applies the lens to whatever range was asked
- Multi-repo rollup applies the lens per repo, then summarises across repos
- Drill-down requests (“why did lead time spike?”) use the supporting metrics in section 2
- Agent reviewing a sprint uses the thresholds to call out what needs attention
Edge cases
- No deployments detected → don’t show zero deploy frequency without a note. Most likely the release mechanism is tag-based or branch-based and not picked up — say so.
- No PRs merged this period → still show commit frequency. Skip PR metrics with N/A.
- Multi-repo comparison → keep one table per repo, do not average across repos. Different teams have different baselines.
- Custom date range that crosses week boundaries → break into weekly rows, not one giant aggregate.
