Why work takes longer than you think
People underestimate how long their own work will take. This is one of the most replicated findings in decision research, it has a name, and it has a documented fix that almost nobody applies.
The interesting part is not that we are bad at estimating. It is that the fix does not require becoming better at estimating.
The planning fallacy
Daniel Kahneman and Amos Tversky named the planning fallacy in 1979: the tendency to underestimate how long a task will take, even when you know that similar tasks have run over in the past.
The clearest demonstration comes from Roger Buehler, Dale Griffin and Michael Ross in 1994. They asked students to predict when they would finish their thesis. The average prediction was 33.9 days. The average actual was 55.5 days. Only around thirty percent finished by the date they themselves had named.
The detail worth sitting with is what happened to the pessimistic estimates. The students were also asked for a worst-case scenario, the date by which they would finish if everything went wrong. On average, the real completion date was later than that worst case. Their pessimism was not pessimistic enough.
Why experience alone does not fix it
The obvious reply is that people learn. You miss a deadline, you adjust, you estimate better next time.
The same body of research says that is not what happens. Knowing that your past projects ran over does not, by itself, correct the next estimate. People treat the current task as different: this time the requirements are clearer, this time there will be no interruptions. The past is acknowledged and then set aside.
What does change the estimate is being made to apply the past to the specific task in front of you. Not remembering it in general terms. Retrieving what comparable work actually took, and using that number as the starting point.
Reference-class forecasting
That method has a name too. Reference-class forecasting means estimating a task by finding a class of similar completed tasks and starting from their actual outcomes, rather than from your mental model of how this one will go.
Bent Flyvbjerg has applied this to large projects at scale. In a database of more than sixteen thousand projects, 8.5 percent delivered on both time and budget. Just 0.5 percent delivered on time, on budget, and with the benefits that had been promised.
The method is not fringe. The UK Treasury's Green Book, the guidance governing how British public spending is appraised, requires explicit adjustments for optimism bias based on evidence from comparable past projects.
The part that stops people
Reference-class forecasting needs a reference class. You need to know what comparable work actually took, not what you remember it taking.
Most people do not have that. Timesheets record what was billed, which is not the same thing. Project tools record when a ticket was opened and closed, which includes every day it sat waiting. App and website trackers record that you spent eleven hours in a browser, which does not say which work those hours went to.
So the ingredient the research says you need is the ingredient nobody has. The failure is not discipline. It is that reconstructing what last quarter's similar project actually consumed is genuinely hard work, and it comes due at exactly the moment you are trying to plan the next one.
What an automatic record changes
If something has been recording what you worked on, in named pieces of work rather than applications, then the reference class already exists. You are not reconstructing it, you are querying it.
That is the specific reason a task-level record is useful for planning where an app-level one is not. "Eleven hours in the browser" is not a reference class. "The last three client onboardings took 14, 19 and 16 hours of active time" is.

In practice the useful question is narrow. Not "how long will this take", which invites a guess, but "what did the last few comparable pieces of work actually take". The first is a prediction. The second is a lookup, and the research says the lookup is the part that works.
Where this stops working
Three honest limits, because a method oversold is a method that gets abandoned.
The estimate is only as good as the match between the reference class and the new task. If you group work that is not really comparable, you get a confident number built on the wrong sample. Choosing the class is a judgement call and the data cannot make it for you.
Anchoring defeats the whole exercise. If you decide the answer is two weeks and then go looking for evidence, you will find some. The record has to be consulted before the estimate is formed, not after.
And some underestimation is not a cognitive error at all. When a low number is what gets a project approved or a bid accepted, the incentive is producing the estimate, and better historical data will not change it. That is a different problem and tooling does not solve it.
See what your own work actually takes
Timeslicer records your day as named tasks automatically, on Mac and Windows, and exposes that history to your AI tools over a read-only interface.
Sources
Kahneman and Tversky introduced the planning fallacy in 1979. The thesis-completion study is Buehler, Griffin and Ross, 1994, which is also the source for the finding that knowing your history does not correct an estimate unless you are made to apply it. The sixteen-thousand-project figures are from Bent Flyvbjerg's research on large project delivery. The optimism-bias requirement appears in the UK Treasury Green Book.
Related reading: how to see what you actually worked on and how automatic trackers compare.
Last updated 3 September 2026.