Hey there! 👋
Welcome back to SavvyMonk, your one-stop for AI and tech news that actually matters.
Today's story is the kind that sounds like a joke until you sit with it. Amazon built a scoreboard to get its engineers excited about AI, the engineers got excited about the scoreboard instead, and the company ended up paying real money to watch them game it. Then a senior leader had to step in and tell thousands of developers to stop using AI for the sake of using AI.
Let's get into it.
Join 2M+ Professionals Getting Ahead on AI
Keeping up with AI shouldn't feel like a second job.
But between the new tools, viral posts, and endless hot takes, most people spend hours every week trying to figure out what actually matters.
The Rundown AI fixes that.
It's a free newsletter that gives you the AI news, tools, and tutorials you actually need to know. All in just 5 minutes a day.
Over 2M professionals at companies like Apple, Google, and NASA already read it every morning to stay ahead.
Plus, if you complete the quiz after signing up, they'll recommend the best tools, guides, and courses for your specific job and needs.
TODAY'S DEEP DIVE
How Amazon's AI Usage Contest Backfired
Earlier this year a group of engineers at Amazon built an internal dashboard they called KiroRank. It sat on top of Kiro, Amazon's in-house AI coding tool, and the concept was simple. It measured how much each person used AI, scored them on that usage, and ranked everyone against everyone else.
On paper the goal was reasonable, since the dashboard was meant to show the rest of the company how much AI could speed up ordinary work and to nudge more people into actually trying the tools. It also fit neatly into a bigger push, because Amazon had already set an aggressive target that more than 80% of its developers should use AI tools every single week, and a public ranking of who used the most slotted right into that mandate.

Amazon Spheres, Seattle, USA | Photo by Alexandra Tran on Unsplash
What Amazon got instead was a clean demonstration of what happens when you turn a measurement into a target. The moment people could see their rank, the incentive stopped being good work and became a higher number, and a workforce full of very smart engineers is exactly the wrong audience to hand a gameable metric.
Goodhart's Law in Action
Once the rankings were visible, the smart play was obvious. The metric rewarded volume, so people manufactured volume. Engineers began pointing AI agents at trivial, low-value tasks and running the same kinds of calls over and over, not because the work needed doing but because every call nudged their score higher.
Silicon Valley already has a name for this behavior, tokenmaxxing, and Amazon's own internal term for it was the same. It spread far enough that multiple employees described feeling real pressure to automate pointless work simply to keep pace on the board, which is the opposite of what a productivity tool is supposed to encourage.
The catch is that none of those tokens are free. Every throwaway call ran on Amazon's own cloud and landed as a real line item on the company's compute bill, so the leaderboard was quietly converting a vanity metric into a rising cost. It produced a ranking stuffed with activity that looked productive on a dashboard and accomplished close to nothing in reality, and the bill kept climbing the whole time.
Treadwell Steps In
The person who finally called it was Dave Treadwell, a senior vice president of engineering. In a message to staff earlier this week, he told them to stop using AI just for the sake of using it and to point it instead at solving customer problems, solving business problems, and actual innovation. He was careful to frame the dashboard as something built with good intentions, but he was clear that it had ended up generating extra cost without adding matching value, so the team pulled it.
Amazon's official line afterward leaned on that same framing, describing KiroRank as a beta dashboard that a group of employees had created to raise awareness of how AI can accelerate work, never a formal or approved company tool, and a spokesperson confirmed it had since been deprecated.
That nuance is worth holding onto, because it changes how big a deal this is. This was not a sanctioned company program that collapsed under its own weight. It was an unofficial experiment that worked too well in the wrong direction and got switched off once leadership saw the bill.
The Cost Problem Underneath
The leaderboard mess sits on top of a much larger spending problem, which is why it stung enough to draw a senior VP into a company-wide message. Amazon plans to spend somewhere around $200 billion in 2026, most of it on AI infrastructure, and token-based pricing means usage and cost move together in a way that older software never did.
Goldman Sachs expects token consumption to rise roughly 24 times by 2030 even as the price per token falls more than 90%, because agentic AI burns somewhere between 5 and 30 times more tokens per task than a basic chatbot. That is the trap in one sentence, since usage is growing so much faster than prices are falling that the total bill keeps going up no matter how cheap each token gets. A leaderboard that explicitly rewards more usage is pouring fuel directly onto that fire.
Amazon Is Not Alone
The same pattern has been showing up across big tech, which is what makes this more than an Amazon quirk. Meta ran an internal leaderboard called Claudeonomics that tracked the token usage of its 85,000 employees, singled out the top 250, and handed out badges with names like Token Legend and Cache Wizard, logging 60 trillion tokens in a single 30-day window before the company shut it down after the data leaked.

Microsoft Office, Hyderabad, India | Photo by Mahesh Bharadwaj on Unsplash
Microsoft recently cancelled the Claude Code licenses its own engineers had fallen in love with, once the bills climbed faster than anyone had budgeted for. And Nvidia, the company selling the chips that power all of this, has admitted that compute now costs its teams more than the engineers who use it. When the chip vendor, the model buyer, and the platform builder are all saying the same thing, it stops looking like a coincidence and starts looking like the actual economics of the moment.
What Replaced It
Amazon did not abandon measurement, it changed what it measures, and that shift is the most useful part of the whole episode. The new internal metric is called normalized deployments, which tracks AI-generated code that actually ships and proves useful rather than raw call volume. The wording change is small and the intent behind it is large, because the goal moved from more AI usage to better AI usage. It is a quiet admission that the first version measured the wrong thing entirely.
The Bottom Line
Counting how much people use AI tells you almost nothing about whether it is helping, and Amazon just learned that its own scoreboard was paying engineers to waste compute on work nobody needed. Pulling it was the right call, and the pivot to measuring shipped, useful output is the part other companies should copy. If you run a team, the lesson is blunt. Reward outcomes rather than activity, because whatever you put on the board is exactly what people will optimize for, right down to the cloud bill.
AI PROMPT OF THE DAY
Category: Process Improvement
"I want to measure how my team uses [AI tool] without rewarding busywork. Design 3 metrics that track real outcomes instead of raw usage. For each one, explain what it measures, how I would collect the data, and one specific way an employee could game it so I can build in a guardrail against that."
ONE LAST THING
Every number you put on a screen quietly becomes a target, whether you meant it that way or not, and people will reshape their work to move it. The trick is to measure the thing you actually want rather than the thing that happens to be easy to count, and Amazon just paid a real compute bill to relearn that lesson in public. The fix was not a fancier dashboard but a better question about what good work actually looks like. So it is worth asking what your own team would start doing differently the day you changed the number on the board. Hit reply, I read every response.
See you in the next one.
— Vivek
P.S. Know a founder, CTO, or engineering lead trying to roll out AI without torching the budget? Forward this their way. They can subscribe at https://savvymonk.beehiiv.com/



