Claude Code in September: the tool finally lets you measure your setup.

- claude-code
- ai
- plugin-eval
- skill-doctor
- coding-agents
- dev-workflow
- anthropic
- productivity
It is Thursday night. You have Claude Code open, a CLAUDE.md that crossed a hundred lines a while ago, a few skills installed, and a couple of plugins you added a few weeks back.
None of it is there by accident. You built it up on purpose, testing what paid off and what did not. And it works.
But stop for two seconds and ask yourself how much each of those pieces actually helps you today, and you do not have a number. You have a hunch, which is a different thing.
For a long time that was the only option: you tuned the setup by feel and trusted it was helping, because measuring it precisely was awkward or flat out impossible.
What Claude Code shipped in September changes exactly that. Yes, there is a bigger model. But the thing that moves your day is different: you can finally measure what you used to tune by eye, and confirm with data what you already suspected.
What September shipped
To see why this matters, it helps to look at what arrived first, without reciting the whole changelog. If you want the week-by-week detail, it is in the official Claude Code What's new.
The week of August 31 to September 4 brought Fable 5.1 into Claude Code. The context window is still one million tokens, the same it has been for months across the models. And that room is part of the problem: everything fits, so it is easy to pile up rules, skills and files without anyone measuring what each one adds.
That same batch added /skill-doctor, which shows you how much context each skill costs you and how often it really gets used. And /diff started opening a live panel beside the conversation that refreshes as Claude edits.
The following week, September 7 to 11, brought the good part: claude plugin eval. It runs your plugin against a suite of test cases, scores the results, and compares them against a no-plugin baseline. claude plugin eval init even drafts the cases and graders so you can get going.
Those same days added maxEffortLevel, a setting that caps the effort level across every provider, and the option to pop any Desktop pane out into its own window.
Put side by side, these features do not look like they have much in common. But they all point the same way: to stop working blind on your own configuration.
From hunch to evidence
Here is the real shift, and it is more about habit than features.
Some context to size it up: according to the JetBrains 2026 survey, with more than 15,000 responses, 90% of professional developers use coding agents at least weekly and 68% use them daily. When you open a tool every day, a bad configuration stops being an occasional annoyance and turns into a tax you pay in every session.
You used to tune it by judgment and move on. Now you can tune, measure, and only then decide, with the number in front of you.
/skill-doctor is the clearest case. You open it and find that a skill you installed a month ago has been used zero times, and it has been eating context on every startup anyway. You remove it. You just made room for what does pull its weight.
plugin eval goes one step further. Say you have a plugin that, in theory, makes Claude follow your repo conventions. Does it really do that better than without it? You run the eval against the baseline and look at the score. If it moves nothing, it was decoration.
A small example
Say you installed a plugin so Claude writes your commits in conventional commit format.
Before, the only way to know if it worked was to read commits for a week and trust your memory. Now you write four or five cases with plugin eval init, define what counts as a good commit, and run the eval with the plugin and without it.
The number tells you whether it helps, whether it makes no difference, or whether it is messing up your output. Ten minutes against a week of gut feeling.
The same goes for effort. If you had everything cranked to the max "just in case", maxEffortLevel forces the uncomfortable question: does this task need the model thinking hard, or am I paying a premium for nothing? Plenty of times a small refactor comes out just as well at half the effort, and you only find that out once you cap it and compare.
Where I still do not trust it
Being able to measure is not the same as understanding, so it is worth easing off the gas.
A green eval measures you against the cases you wrote. If your cases lean toward what you already expected, so will your confidence. The score is only as good as the questions you asked it.
/skill-doctor tells you what got used and what it cost, not whether it got used well. A heavily used skill might be jumping in where you never called it, and the number still shows green.
And more effort does not always mean a better result. Sometimes it is the same output, slower and pricier. That is why the ceiling is as useful as the floor: it forces you to think about how much you actually need for each thing.
Measuring your setup is a new tool. It is not an autopilot that excuses you from looking: it gives you numbers, and reading them is still your job.
The interesting part of finally being able to measure your setup is not that it corrects you. It is that for the first time you can put a number on what you used to run from memory.
The tool became measurable. Now it is on us to build the habit of looking.
How much of your AI setup have you ever reviewed with a number in front of you, and how much still runs on hunch alone?