Technical performance answers only part of the question

A technology may meet its technical objective and still fail to address the needs of the people expected to use it. Measuring speed, accuracy, cost, or consumption helps describe performance under specified conditions; it does not, by itself, explain who defined those conditions, which problems were left out, or how benefits and burdens are distributed. Social evaluation does not replace technical testing: it complements it with questions about the context and consequences of innovation.

The distinction matters because a metric is not a neutral description of a system’s entire value. It is a selection of what has been chosen for observation. If, for example, a tool is designed to reduce the time required for a task, it is worth asking which task is made faster, for whom, and whether the time saved actually becomes an improvement valued by the people affected. The useful question is not only whether it works, but for whom it works, under what conditions, and with what possible side effects. The answers require case-specific evidence; they cannot be inferred from advertised capabilities.

What a metric leaves out

An indicator can describe the aspect for which it was designed precisely while failing to measure other important outcomes. Before comparing figures, it is therefore necessary to know what they represent, how they were obtained, and which conditions were held constant. A performance measure answers a bounded question; it does not replace a broader assessment of experience, access, or the distribution of consequences. This distinction prevents an improvement in one indicator from being treated automatically as proof that the innovation as a whole is beneficial.

Who participates and which forms of knowledge are recognised

Understanding an innovation means tracing the decisions that make it possible: who identifies the problem, who funds development, who sets requirements, which groups participate in design, and who is responsible for deployment. It is also worth asking who can express a need and whether that contribution changes specific decisions. Nominal participation—a consultation with no demonstrable influence—is not the same as sharing decision-making power.

Science, technology and society studies help frame this contextual perspective. A text from the Organization of Ibero-American States discusses the relationship between technology, society and education, and presents an approach that incorporates technical and cultural dimensions, rather than focusing only on teaching tools. That perspective can help formulate questions, but it does not, by itself, demonstrate the impact of any particular contemporary technology. An analytical framework guides research; it does not replace case-specific data. (The source is a conceptual reflection published in an education journal, not an experimental evaluation of a product.)

It also matters which forms of knowledge are considered valid. Laboratory data, professional experience, and the practical knowledge of people who use a service can describe different aspects. The point is not to choose one and discard the others, but to make clear what each contributes, how it was obtained, and what limitations it has. In practice, this may mean documenting disagreements between users and technical decision-makers rather than reducing them to a single average figure. Making such differences visible helps establish whether they reflect different needs, different conditions of use, or different definitions of success. It does not resolve disagreement on its own, but it prevents that disagreement from being hidden behind an aggregate measure.

What evidence is needed to evaluate a case

Before attributing an effect to a technology, the system being evaluated and its context must be defined precisely: version, population, task, setting, and observation period. The next step is to identify the outcome to be explained and the comparison being used. Without these details, two reports about the same type of tool may concern populations or conditions so different that their results cannot be compared directly.

Evidence can combine technical documentation, observation of use, interviews, administrative records, and outcome analysis, provided the method fits the question. A study of educational practices mediated by digital technologies, published in RIED, provides an example of research focused on evidence of learning in an educational context. Its subject does not allow results to be extrapolated to other sectors or to every technology: it shows why an evaluation should specify the practice and outcome rather than treating “the digital” as a uniform intervention.

The choice of methods also determines what can be known. Technical documentation can describe the system; observation and interviews can provide information about how it is used and experienced; outcome analysis can help examine changes associated with the intervention. None of these sources necessarily answers every question. Each claim in a report should therefore be linked to the evidence supporting it, with any aspects not covered by the available data identified. This traceability helps distinguish a conclusion supported by the study from a broader interpretation that would require additional evidence. A large volume of information does not remove the need for that information to be relevant to the question.

A practical reading protocol

  1. Define the case: product or system, version, task, location, and dates.
  2. Identify the people affected: direct users, workers, decision-makers, and those who might be excluded.
  3. Check the method: sample, comparison, measures, missing data, and declared conflicts of interest.
  4. Separate outcomes: technical performance, user experience, distribution of effects, and unintended consequences.
  5. Look for limitations and replication: compare relevant studies without assuming that a result automatically generalises.

Observed results are not guaranteed impacts

A conclusion must match the research design. If a study observes an association, that is not enough to claim that the technology caused it; other changes may have occurred at the same time, or there may be differences between people who use the system and those who do not. Even when there is an appropriate comparison, it remains necessary to check whether the population and setting resemble the case to which the result is meant to apply. Generalisation is an inference, not a directly observed fact.

A systematic review of telemedicine interventions for people with multimorbidity in primary care illustrates, through its scope alone, why population, intervention, and health outcomes need to be specified. It is not general evidence about all telemedicine or every digital innovation. To assess its conclusions, the full article must be consulted and its methods and results reviewed; the title or available abstract does not justify assigning it an effect size here. The difference between the evidence available and the evidence it would be desirable to have should be made explicit.

For that reason, reports should distinguish three levels: what was observed, what interpretation the research team proposes, and what is expected to happen at a larger scale. Projections may be reasonable, but should not be presented as impacts already demonstrated. Similarly, a lack of data about a group does not prove that the group was harmed; it identifies a gap that limits what can be claimed. Keeping these levels separate helps readers assess how much a conclusion depends on observed data and how much on interpretation or expectation. It also helps specify what further information would be needed to strengthen, revise, or narrow the conclusion, without turning uncertainty into a claim the evidence cannot support.

How to check the scope before accepting a conclusion

Critical reading starts with the primary source: the relevant article, evaluation report, protocol, or official documentation. Summaries and explanatory articles can help locate material, but may omit methodological decisions. Check authorship, date, population, the definition of the technology, and the relationship between funding and evaluation. If the document links to data or materials, their availability makes scrutiny easier, though it does not by itself guarantee a sound method.

Next, check whether the published claim preserves the scope of the original. A conclusion about one school, health centre, or specific group is not a universal rule. Look also for null results, different effects across subgroups, and limitations acknowledged by the authors themselves. If these details cannot be verified, the responsible approach is to make a narrower statement: describe what is known and identify what remains unresolved, rather than filling the gap with a prediction.

Evaluating an innovation means connecting capabilities, decisions, and experiences with evidence proportionate to each claim. This does not mean demanding absolute certainty before adopting anything. It means making success criteria visible, listening to those involved, examining who bears the costs, and revisiting conclusions as new information emerges. Without a specific case and sufficient primary sources, it is not possible to issue a verdict on particular impacts; it is possible to establish a method that makes such a verdict verifiable and limited to what was actually investigated. A carefully bounded conclusion is no less useful for acknowledging its limits: it clarifies what the evidence supports, what remains open, and which questions a later evaluation should address.