Principles for Understanding Trust in Artificial Intelligence: A Conceptual Map and the Research Questions Left Open

Citation: Everett, J. A. C., Claessens, S., Knöchel, T.-D., & Reinecke, M. G. (2026). Principles for understanding trust in artificial intelligence. Nature Reviews Psychology. https://doi.org/10.1038/s44159-026-00562-1

1. Overview of the Article

The review by Everett and colleagues, published in Nature Reviews Psychology, takes a deliberate departure from the dominant convention in the field. Instead of compressing a decade of empirical findings on trust in artificial intelligence (AI) into yet another meta-analytic summary, the authors propose six theoretical principles intended to structure a rapidly expanding and increasingly fragmented literature. Their central diagnosis is that the core problem of the field is not data scarcity, but rather the conceptual ambiguity of the term “trust” itself: a multilayered construct, frequently conflated with adjacent notions, and lacking a shared interdisciplinary definition. The review therefore functions less as an empirical synthesis and more as a conceptual map of the field.

The animating argument is that trust in AI is not a single attitude or variable, but an inferred, multidimensional, dynamic and context-sensitive process. “Trust in AI” is not one thing; it varies across systems, individuals and contexts. The authors adapt classical interpersonal trust models from psychology — particularly Mayer, Davis and Schoorman’s 1995 organisational trust model — to the AI context, while explicitly acknowledging the conceptual difficulties this adaptation entails.

The theoretical stance of the authors goes beyond the operational question of “should trust be increased?” to foreground a more demanding normative question: “under which conditions, in which systems, and for which purposes is trust normatively defensible?” In this respect the review reads less as a conventional literature survey and more as a measured intervention into the ethical-political positioning of the field.

2. Summary of the Six Proposed Principles

The authors propose six interrelated yet distinguishable principles to structure the field. These are presented not as rigid rules but as a heuristic framework for guiding future inquiry.

  1. Trust in AI is inferred. Trustworthiness is not an objective property that can be “engineered into” the machine; it is a quality attributed by the user through ordinary processes of social impression formation. Actual trustworthiness and perceived trustworthiness are distinct, as evidenced by the inconsistent effects of explainable AI (XAI) on trust outcomes.
  2. Trustworthiness, trust and trusting behaviour are distinct. Judging a system to be trustworthy, holding an attitude of trust toward it, and actually relying on it in practice are conceptually and empirically separable. Users may declare an AI untrustworthy yet behaviourally comply with its recommendations, or the reverse may hold.
  3. Trust rests on both performance and morality. Whereas the classical literature predominantly framed trust along the performance axis (competence, reliability), contemporary generative AI systems make the moral dimension of trust (sincerity, ethical alignment) equally salient. A system can be highly capable but deceptive, or ethically aligned but unreliable.
  4. Trust in AI is agent-specific. The question “do people trust AI?” is misplaced; the proper formulation is “trust which AI to do what?” Users do not extend the same evaluative criteria to GPS navigation systems, generative chatbots and autonomous weapons systems.
  5. Trust in AI is individually variable. Personality traits, AI literacy, attachment style, age, gender and — most consistently — cultural background shape both the level and the bases of trust. The same system may receive cautious adoption in one cultural context and outright rejection in another.
  6. Trust in AI is strategically motivated. The same individual may trust the same AI to different degrees in different contexts, depending on personal incentives, motives and the desire to deflect responsibility. This is the least studied of the six principles and, simultaneously, the most consequential for applied domains such as patient safety.

3. Relevance for Health Management

From the standpoint of health management, this review offers more than a theoretical contribution; it carries direct operational implications. From clinical decision support systems to patient triage, from radiological image interpretation to drug-interaction alerting, a broad swathe of healthcare delivery is being restructured around AI. The success of this transformation depends on the extent to which clinicians, nurses and patients trust these systems, on what grounds, and under what conditions.

The trustworthiness–trust–trusting behaviour distinction proposed by the authors illuminates a paradox routinely observed in healthcare organisations: clinical staff may verbally endorse a decision support system as trustworthy without actually complying with its recommendations, while in conditions of time pressure they may comply with systems they have explicitly criticised as untrustworthy. This finding indicates that clinical adoption failures cannot be reduced to a problem of “insufficient training”; they are structural.

The principle of strategic motivation is particularly critical for healthcare. In contexts where responsibility can be “laundered” onto the AI (agency laundering), the risk of clinical reasoning atrophy and of overtrust patterns culminating in adverse events increases. The systematic empirical investigation of this phenomenon in healthcare contexts is overdue.

4. Research Questions That Users of This Article Cannot Directly Answer

The review offers a conceptual map; the researcher who attempts to use it for empirical work will rapidly notice that many operationally indispensable questions remain unresolved. The questions below represent gaps that the framework itself opens but does not close.

4.1. Measurement and Operationalisation Gaps

In Table 1 the authors compare eleven scales used to measure trust in AI and implicitly suggest that none of them captures the full set of six principles. They do not, however, propose a concrete measurement strategy that follows from this diagnosis. The applied researcher is therefore left with the following open questions:

  • How can a psychometrically defensible, multi-system scale that empirically distinguishes performance trust from moral trust, and whose cross-cultural measurement invariance has been formally tested, be developed?
  • If the principle of agent-specificity requires that scales be re-adapted for every system, how can cross-system comparability be preserved? Or should the loss of comparability be accepted as a deliberate methodological choice?
  • Should the divergence between attitudinal and behavioural trust be modelled through classical attitude–behaviour inconsistency paradigms, or does it require a new theoretical framework specific to human–AI interaction?
  • Given the authors’ acknowledgement that the canonical trust game paradigm does not validly transfer to AI (because machines do not value money in the way humans do), what alternative incentivised paradigms can be designed to capture behavioural trust?

4.2. Gaps Concerning the Interaction Between Dimensions

The review defines the six principles separately but does not empirically clarify how they interact. The following operational research questions remain open:

  • In which systems do performance trust and moral trust converge, and in which do they diverge? When a system is high in performance but low in moral trust (autonomous weapons being the canonical case) or low in performance but high in moral trust (an ethically designed but inaccurate diagnostic AI), which dimension dominates the user’s adoption decision?
  • How do individual differences (personality, culture, AI literacy) interact with agent-specificity? Does high extraversion increase trust toward all classes of AI, or only toward socially situated systems?
  • Which of the two evaluative dimensions — performance or morality — is more elastic under strategic motivation? When the AI’s output is congruent with the user’s interest, is the system perceived as more competent or as more ethical?

4.3. Longitudinal and Temporal Validity Gaps

The authors explicitly invoke the “temporal validity problem” and warn that current findings may have an “expiration date”. This warning, however, leaves the researcher without a methodological blueprint:

  • How should a longitudinal design be constructed to track the evolution of trust across versions of a single AI system (for instance, GPT-3.5 → GPT-4 → GPT-5)?
  • Does the relationship between initial trust and sustained trust in AI follow a different pattern than in human–human relationships, and if so, in what respects?
  • Following an adverse event (a diagnostic error, a fatal autonomous-vehicle collision), through which mechanisms is trust recovered? Is the erosion of trust attributed primarily to the system itself, to its developer, or to the broader category (“AI in general”)?

4.4. Gaps Concerning Strategic Trust and Agency Laundering

The authors emphasise that the sixth principle is the least empirically developed dimension. It is also the one most directly relevant to applied health management research:

  • Do clinicians comply more with AI recommendations in clinical decisions where responsibility can be plausibly shared, compared with decisions made in isolation? How does this pattern correlate with adverse-event rates?
  • Does the “sycophantic AI” pattern (systems that confirm the user’s prior position) also manifest in clinical decision support systems, and if so, how does it affect the calibration of clinical judgment?
  • Does strategic trust operate at the institutional level as well? Are hospital administrations inclined to extend miscalibrated trust to AI systems onto which financial or legal liability can be transferred?

4.5. Normative and Ethical-Political Gaps

In the closing section the authors clearly insist that the goal of “more trust” should not be accepted uncritically; yet they do not equip the researcher with the tools needed to translate this critique into empirical work:

  • How is calibrated trust to be measured by an observer who does not have access to the system’s true level of trustworthiness? Who, with what authority, sets the reference point for calibration?
  • How can the empirical claims of “ethics-washing” and “machinewashing” be tested? Through which experimental designs can the hypothesis that certification labels generate unwarranted trust be evaluated?
  • How can the algocracy critique be operationalised within specific sectors such as health management? Through which indicators is the atrophy of clinicians’ ethical skills (“ethical deskilling”) to be tracked?

4.6. Gaps Specific to Türkiye and Comparable Contexts

The review acknowledges that cultural differences shape both the level and the bases of trust, but it does not offer a dedicated analysis of countries — such as Türkiye — that combine an emerging-economy profile with a long-established institutional health system. The questions that remain open for the health management researcher in this context include:

  • To what extent do the bases of clinical AI trust among physicians, nurses and patients in Türkiye diverge from the patterns reported in Western literature? Does the relative weight of the performance and moral dimensions vary across cultures?
  • How do dimensions such as power distance and uncertainty avoidance modulate compliance with AI recommendations in clinical settings?
  • In environments where a language barrier exists (limited Turkish-language interfaces), how should “linguistic fit” be operationalised as an additional dimension shaping user trust?

5. Concluding Assessment

The six principles proposed by Everett and colleagues constitute a powerful conceptual map for organising the literature on trust in AI. Yet the careful reader will notice that the review opens substantially more questions than it closes. The article is most usefully read as a research agenda document: it identifies, in a disciplined way, what the field does not yet know. The gaps it leaves open — from scale development to longitudinal designs, from the experimental investigation of strategic motivation to cross-cultural validity and ethical-political calibration — sketch the empirical agenda for at least the coming decade.

For the health management researcher, perhaps the most important takeaway is this: the question of “trust in AI” articulates directly with clinical adoption, patient safety and organisational accountability structures, and this articulation has not yet been empirically mapped. The review provides a robust conceptual foundation for that mapping; the construction itself is left to its readers.

Subscribe to the Health Topics Newsletter!

Google reCaptcha: Invalid site key.