Speed alone is not a sufficient measure
The promise of generative AI at work is often expressed in terms of time: drafting sooner, summarising more documents or preparing a response in fewer steps. These may be useful outcomes, but saving time on one task does not automatically mean creating more value across the work as a whole. To support that conclusion, we would need to know which task was measured, what criteria were used to assess the result and what happened to the time saved.
A practical assessment should look at least at the volume of work completed and its quality. It should also account for corrections, human review, errors and new tasks that arise when the tool is introduced. If a response is generated faster but requires extensive checking, the net saving may be smaller than the gross time saved. And if more is produced at the expense of quality or autonomy, a single metric will not adequately describe the effect.
From time saved to the overall result
Time freed up represents a gain only if it can be put towards other useful tasks without shifting additional work elsewhere in the process. A clear comparison should therefore consider the entire cycle: from the moment a task begins until its output has been reviewed and is ready to use. Looking only at the point of generation leaves out steps that also consume resources and may change how the benefit should be assessed.
Widespread use does not prove a productivity increase
A survey reported by the European Commission in October 2025 found that one in three workers in the European Union used AI tools at work. The same communication describes experiences as mostly positive, while also noting that AI-based monitoring and management technologies can increase stress and reduce autonomy. These are relevant indications for understanding adoption and people’s experience; on their own, however, they are not a measurement of productivity.
The distinction matters because a survey about use or perceptions answers different questions from an experiment comparing work outcomes with and without assistance. The former can help describe who uses tools and how they assess the experience. To attribute a productivity improvement to AI, an explicit definition of the outcome is also needed, along with a comparison that can separate the tool’s effect from factors such as experience, task or work organisation. The publicly available summary should not be presented as though it settled these questions.
The proportion of users tells us about the presence of these tools at work, not how much they change the result of each task or what the balance is between benefits and costs. Likewise, a positive experience describes an assessment, but does not replace a measure of quality, quantity or total time. Keeping these questions separate makes it possible to recognise what the survey contributes without assigning it conclusions it does not offer.
What the ILO’s occupational studies contribute
The International Labour Organization has published analyses of the potential effects of generative AI on the quantity and quality of employment, as well as a refined index of occupational exposure. These studies are relevant to framing the debate: they help examine which occupations or tasks could be affected. Exposure to a technology, however, does not mean that an occupation will disappear or that its productivity will increase.
This distinction also helps in interpreting headlines and announcements. An exposure map can indicate where a technology might have a role, but it does not demonstrate that the technology is used in a real-world situation, works well for a particular task, or delivers benefits that outweigh its costs and risks. The ILO approaches the phenomenon through employment and tasks; that kind of analysis should not be turned into a business-performance figure that the source does not provide.
In other words, exposure is a way to define what is worth studying, not the outcome of a performance evaluation. Between the possibility that a tool could play a role in a task and a proven improvement lie questions about use, quality and effects on work that require specific analysis. This avoids treating an indicator about occupations as equivalent to a direct productivity measurement.
How to assess a specific claim
Before accepting a productivity figure, it is worth checking whether the publication explains who took part, what work they did, how long the evaluation lasted and how the result was assessed. It is also important to ask whether the comparison covers ordinary tasks or narrowly defined exercises, and whether people judged the outcome using stated criteria. Without that information, it is difficult to know whether an improvement would hold outside the setting that was evaluated.
A short checklist can help avoid confusing a striking demonstration with evidence that can be generalised:
- Outcome: Is the measure time, quantity, quality or a combination? Does the study explain how quality is scored?
- Comparison: Is there a reference group or condition without the tool?
- Total cost: Are review, correction, training and workflow changes included?
- Scope: Are the population, task type and duration specified?
- Independence: Who funds or publishes the study, and can its methods and limitations be consulted?
What can be concluded, and what remains open
The cited evidence allows us to say that AI tools are already part of the work of a considerable share of the population surveyed in the European Union, and that workplace experiences are not uniformly positive or negative: alongside positive impressions, the Commission points to risks associated with monitoring and management. The ILO’s analyses provide a framework for examining exposure among tasks and occupations. As summarised in the available sources, none of these findings on its own demonstrates a general, net improvement in labour productivity.
For an individual or team considering these tools, the reasonable course is to measure a specific, limited use case: compare equivalent tasks, decide in advance what counts as quality and record the full time involved, including review. That exercise can inform a local decision, but it would not prove that the result applies to other sectors or companies. The proportionate conclusion is therefore twofold: there is evidence of adoption and reason to study workplace effects; the scale and reach of any productivity gain require more specific and transparent evaluations.
A local test will be more informative if tasks remain comparable and the approach to assessing the result is decided in advance. Recording the full time involved makes it possible to observe both the work done with the tool and the subsequent review, while an agreed quality standard prevents the output from being judged on speed alone. Even so, the test is limited to the context observed: it does not automatically establish what would happen with other tasks, teams or companies.