Series B fintech, ~200 engineers. Anthropic, OpenAI and AWS. Forty-one days from first invoice to verified reduction. No change to model output quality, no change to system performance.
What follows is how that number was found. Same spend, same models, two ways of looking at it — only one of which can be acted on.
01 / Before · Vendor telemetry
This is the complete picture available from the platform's own usage and cost view. Totals are accurate and real-time. Attribution stops at the API key, and four of the seven keys were shared across teams.
Which team spent the $31,880?
Key shared by platform, QE and data science.
What did we get for it?
No link to tickets, PRs or tests.
How many agents are running?
Agent traffic is indistinguishable from human.
The gap is not a defect in the console. It is a billing system, and it bills well. Cost data lives in one platform and delivery data lives in another, and nobody had joined them. That join is the entire engagement.
02 / After · Jira + eazyBI
Rebuilt inside the client's own Jira, using eazyBI REST source cubes against APIs they already owned. Every figure below traces to a person, a squad, an epic and a delivery artifact — and every panel carries a recommendation rather than only an observation.
03 / Quality as the control variable
Cost measures alone will reward a team for producing less. Every efficiency figure is therefore read alongside the quality signal for the same sprint, drawn from Xray, SonarQube, Snyk and Jenkins. A cost reduction that moves any of them the wrong way is reported as a regression.
Every dashboard in a cost engagement should be capable of embarrassing the engagement. If nothing on it can ever look bad, it is marketing rather than measurement.
04 / How the join works
Nothing sits in the delivery path. Every connector is a scheduled read against an API the client already owned, normalized to the Jira issue key and the commit SHA — the two identifiers every one of these tools already carries.
Attribute the spend. Usage rows carry an api_key_id; a managed reference table in Jira maps each key to a person, squad and cost centre.
Attach it to output. Branch names, build stamps and Xray executions all carry the Jira key, so cost lands on a ticket rather than on a token count.
Divide, then compare. Calculated measures produce cost per unit of delivery, compared against that same team's pre-AI baseline.
05 / What moved the number
06 / What kept it down
Once every dollar resolved to a person, the shape of the org appeared. Nine people accounted for 44% of spend; four of them sat in the bottom quartile for merged PRs, closed tickets and authored tests.
Four ninety-minute working sessions, run against each person's own transcripts with their own spend chart open. Habits changed: scoped context instead of pasted repositories, one well-formed request instead of thirty conversational turns, task class chosen before model.
The same analysis surfaced seven engineers producing above-median output at below-median cost. Given protected time, one per squad, ownership of the CLAUDE.md hierarchy and internal skills after v1, and a fortnightly public transcript clinic in Confluence.
Release notes to shell, dependency and licence reporting to Python against the Artifactory and Snyk APIs, log triage prefiltered by rule before anything reached a model, test fixtures seeded and committed, and a nightly cost-anomaly watch that opens its own Jira ticket.
# the rule the teams now apply
Is the output exact and verifiable? → script it
Does it run more than once a week? → script it
Does it need judgement or language? → model it
Both? → script the filter, model the residue
Configuration produced roughly two-thirds of the reduction. This third is the part that holds after the engagement ends, because it lives in habits and repositories rather than in a settings page. Everything above runs inside the client's infrastructure, on their accounts, owned by them outright — no hosted middle layer, no key held, nothing to disconnect.
After my initial analysis, I give you my estimate on the potential amount I can save you — with 1 free actionable advice you can implement immediately and see the result. Read-only access, scoped to the analysis. Anything we pull is deleted within 30 days.
Book a demo