Productivity
Measuring Your Own Productivity Honestly in the AI Era
By Jim Vernon, Editor, AI Intelligence International · Published 21 February 2026 · Reviewed against our editorial standards · About the author
Most personal productivity metrics were proxies for effort: words written, tickets closed, emails answered. When a machine can inflate every one of them, they stop measuring anything.
This article covers what still carries signal, how to track it without building a second job, and the flattering metrics worth deliberately ignoring.
Key takeaways
- Metrics that stopped working: Volume of anything.
- Metrics that still carry signal: Decisions made and not revisited.
- Tracking without building a second job: One line per week: the most useful thing you did and why it mattered.
- The rework signal: Track how much of your output needed significant correction later.
Metrics that stopped working
Volume of anything. Words, drafts, documents, commits, messages — all trivially inflatable now, and inflating them is often actively harmful to whoever has to read the result.
Hours worked. It was always a weak measure and it is now uncorrelated with contribution in most knowledge roles.
Response speed. Fast replies used to signal engagement; they now signal very little, because speed is free.
Metrics that still carry signal
Decisions made and not revisited. A decision that holds for a quarter represents real work, and this is remarkably hard to fake.
Problems that stopped recurring. Fixing a class of problem is the highest-value activity in most jobs and shows up in no volume metric.
Work someone else built on. If a colleague used your output as a foundation, it was good enough to depend on, which is the strongest available quality signal.
Tracking without building a second job
One line per week: the most useful thing you did and why it mattered. Fifteen weeks of that is a better performance record than any tracking system.
One monthly count: how many things you shipped that a person outside your team noticed. Small, but very hard to inflate.
One quarterly question: what would have happened if you had not been there? Uncomfortable, and the most informative measure available.
The rework signal
Track how much of your output needed significant correction later. Rising rework alongside rising volume is the classic pattern of AI-assisted work going wrong.
Rework is also the number to watch when deciding whether to increase your reliance on generated drafts. Speed that creates correction is not speed.
Ask a colleague rather than estimating it yourself. Self-assessed rework is systematically underestimated by everyone.
Avoiding metrics that flatter
Any metric that improves when you use a tool more is a usage metric, not a productivity metric. Tool vendors report these for obvious reasons.
Any metric with no failure mode is decoration. If a number can only go up, it is not measuring a trade-off, and all real work involves trade-offs.
If a metric has never once told you to stop doing something, replace it.
Reporting this to a manager
Lead with outcomes and their durability, not with tooling. Nobody promotes an efficient user of software; they promote someone whose decisions held.
Be specific about what you verified when using AI. In review-heavy environments, demonstrated judgement is now a differentiator rather than an assumption.
Name one thing that did not work and what you changed. It is the fastest way to make the rest of the account credible.
What to measure instead of hours saved
Hours saved is unverifiable and self-flattering. Three measures are checkable: how long a named recurring task takes end to end, how many of those tasks you complete in a week, and how many needed rework after someone else looked at them.
Pick one recurring task and time it for two weeks before changing anything. Without a baseline, every improvement is a feeling, and feelings about productivity are reliably wrong in the optimistic direction.
Rework is the measure people skip and the one that decides whether a saving is real. A draft produced in four minutes that takes twenty to repair is not a four-minute draft.
The two-week self-experiment
Week one, work as normal and log start and finish times for the chosen task. Week two, use the tool and log the same thing plus review time. Compare medians, not best cases.
Keep everything else constant. Changing tool, process and schedule in the same fortnight guarantees you cannot attribute the result, which is how most personal experiments end inconclusively.
Write the result down even when it is disappointing. A recorded null result stops you re-running the same experiment enthusiastically in three months.
Where the time actually goes
Most knowledge work time is lost to waiting, context switching and unclear requirements rather than to typing. Tools that speed up typing therefore return less than expected, which explains a lot of the gap between vendor claims and lived experience.
Audit a normal day in thirty-minute blocks for three days. The pattern usually shows two or three interruption sources doing more damage than any slow task.
Fix the largest interruption source first. It is almost always cheaper than any subscription and it makes every later tool measurably more effective.
Frequently asked questions
Why are output-volume metrics no longer useful?
Volume is trivially inflatable with generation, and inflating it usually harms whoever reads the result. Measure durable decisions and recurring problems solved instead.
What is the single best personal productivity signal now?
Work that someone else built on. If a colleague depended on your output, it met a quality bar that no volume metric can capture.
How do I know if AI is actually helping my output?
Track rework. Rising volume with rising correction means the speed is being paid back later — and ask a colleague, because self-assessed rework is always underestimated.
Is time tracking worth the overhead?
For a fortnight on one task, yes. Permanent tracking of everything costs more attention than it returns and people quietly abandon it.
What if the numbers show no gain?
That is a useful result. It usually means the bottleneck was elsewhere, not that the tool is bad, so look at requirements and interruptions next.