Why DORA Metrics Are Necessary But Not Sufficient
The four DORA metrics (Deployment Frequency, Lead Time for Changes, Change Failure Rate, and Mean Time to Recovery) have become the gold standard for measuring software delivery performance. And for good reason: they're backed by years of research from the DORA team at Google, validated across thousands of organizations.
But here's the uncomfortable truth: DORA metrics alone can create false narratives.
The problem with static benchmarks
Consider two teams in the same organization:
Team A manages infrastructure with Terraform. They deploy once a month, after CAB approval and thorough change window planning. Zero incidents in six months.
Team B ships a React frontend multiple times a day. No CAB, no change windows, feature flags everywhere. Three incidents this quarter.
According to DORA's universal thresholds, Team A is "Medium" and Team B is "Elite." But is that really the full picture?
Team A might be performing exceptionally given their constraints. Team B might be underperforming given their freedom to ship. Static benchmarks can't tell you which is which.
Constraints are not excuses, they're context
Every team operates under different constraints:
CAB / Change Advisory Boards limit deployment frequency regardless of team capability
Store reviews (App Store, Play Store) create hard deployment ceilings
Regulatory gates add mandatory wait times
Technology constraints: Terraform and embedded systems don't deploy like microservices
These constraints aren't failings. They're the playing field. Judging a team without accounting for their playing field is like comparing a swimmer's time to a runner's.
What contextual benchmarking adds
The idea is simple: compare what you don't control (constraints), improve what you do (practices).
Instead of comparing every team against universal thresholds, group teams by their constraints, then compare within those groups. Two frontend teams with store review constraints? Now that's a fair comparison. If one deploys 4x more than the other, the difference is in practices (feature flags, trunk-based development, PR discipline), not in circumstances.
This is what CodeSpectra's Contextual Benchmarking does. It doesn't replace DORA, it makes DORA actionable by adding the context that static thresholds miss.
Beyond Accelerate's playbook
The Accelerate study identified specific capabilities (feature flags, trunk-based development, test automation) that improve delivery performance. But those capabilities were studied in a pre-AI world.
AI coding tools fundamentally change how code is written and reviewed. The capability playbook from 2019 doesn't fully apply in 2026. Contextual Benchmarking adapts because it learns from what actually works for teams with similar constraints, right now, not from a survey from years ago.
Start with DORA. Don't stop there.
DORA metrics are the foundation. Every engineering organization should measure them. But if you stop at "Elite / High / Medium / Low," you're missing the story behind the numbers.
The real question isn't "what's our DORA rating?" It's "what should we do about it, given our specific constraints?"
That's what contextual benchmarking answers.
CodeSpectra measures DORA metrics automatically from your CI/CD pipelines: GitHub, Azure DevOps, GitLab. Contextual Benchmarking is available on the Pro plan. Start free with up to 30 repos.