1 Introduction

Currently, robots appear in a wide range of scenarios. Since their earlier functions focused on the industrial world, they have increasingly become more human-like, physically and psychologically. Some of them have entered the private spheres of our lives. This is the case with social robots: machines with a physical body — whether human-like, animal-like, or other — that interact with human users. This interaction entails the capacity to recognize and produce language. Not without controversy, these robots might also be able to recognize and emulate human emotions. They are currently used as caretakers or social companions for older adults, as assistants in therapies for children with learning or socializing difficulties, or as sexual and intimacy partners for consumers.

Social robots entail numerous philosophical implications. Their similarity to humans in appearance and their different capacities (granted by sophisticated AI that continues to improve) give rise to questions such as whether they have — or could potentially have — moral agency or moral patience, and therefore possibly responsibilities and rights (Coeckelbergh, 2010b; Gunkel, 2022; Schwitzgebel & Garza, 2015; Müller, 2021). These issues are directly related to how we treat and should treat robots and what consequences may arise if we, for example, rape a sex robot or kick a robot dog (Carpenter, 2017; Darling, 2021). In this paper, we will focus on three specific and unavoidably interrelated questions concerning Human-Robot Interaction (HRI) that are directly linked to how we behave toward robots: (i) can humans trust social robots?; (ii) can social robots be trustworthy?; and (iii) should we trust social robots?

In the emerging literature, there are many references to the notion of “Trustworthy AI.” (TAI). In 2019, the High-Level Expert Group of the EU developed a set of guidelines declaring that AI, in order to be trustworthy, should meet three conditions: it must be lawful, ethical, and robust. However, as Margrit Sutrop (2019) notes, “Although there is much talk about trust [in the guidelines], surprisingly little is said about what constitutes trust and what it depends upon”. Similarly, many authors talk about trust in AI, Artificial Agents (AAs), or robots in a very broad sense, either proposing new definitions such as e-trust (Taddeo, 2010) or not clarifying what they understand for trust (Gillath et al., 2021; Choung et al., 2023; Durante, 2010). Nevertheless, in most of these cases what is being discussed is reliance rather than trust. Trust, in the relevant sense — interpersonal trust — appears in more complex instances than our relationship with, for example, a smartphone. Therefore, some authors argue that reliance should be the term used when we talk about our attitudes towards any type of current technology: the label “Reliable AI” should replace TAI (Ryan, 2020; Fossa, 2019). Other approaches emphasize how we actually treat artificial agents, arguing that there is room for the attitude of trust to be developed towards them (Sweeney, 2022; Coeckelbergh, 2012b). Still, a clear distinction between trust and reliance is needed to determine when each concept is accurate in philosophical and legal literature. To do so, we need to return to deeply philosophically rooted questions, which relate to those posed at the end of the last paragraph: What is the nature of trust? What is trustworthiness? When is it warranted?

In Sect. 2 of this paper, we will investigate how humans treat robots. Can humans develop attachment and empathy towards them? Can they consider robots as friends and even fall in love with them? Different attitudes that humans manifest towards technology (as we will see, not only social robots) suggest that interpersonal trust already seems to exist in certain contexts of Human-Technology Interaction. In Sect. 3, we will address questions concerning the nature of reliance, trust, and trustworthiness by drawing on the traditional philosophical literature on this subject. In Sect. 4, we will review the literature that has reflected on our attitudes towards AI and robots, sometimes specifically concerning trust and sometimes their moral standing. Our own view encapsulates an apparent paradox: there is trust without trustworthiness. We will synthesize some of the previously mentioned views, noting that while artificial agents cannot be considered trustworthy — at least at the present stage of technology — humans will still trust them. In the final section, we will ask whether we should trust social robots. This question is important because our attitudes of trust towards the non-trustworthy will have ethical implications.

2 How do we Treat Robots?

2.1 Anthropomorphism: Emotional Attitudes towards Robots

There’s little reason to believe we won’t develop relationships with robots, anthropomorphize many of them, and treat some of them like our companions. As robots enter into our homes and lives, we are almost certainly going to bond with them (Darling, 2021).

Humans have a biological tendency to anthropomorphize. From drawing faces on rocks to treating animals as if they understood philosophical concepts, we ascribe human attributes to objects and other species — attributes that they lack. For the most part, we are aware that it is a projection, yet we still are inclined to do this. The more the object resembles humanity, whether through appearance (body or face), communication, or movement, the stronger this propensity grows. This disposition naturally arises when we relate to robots, especially social robots, such as sexual ones, since they attempt to be as human-like as possible in appearance and behavior. In addition, as Carpenter notes, our anthropomorphism is not diminished by our knowledge of the nature of the robot (2017). The development of stronger emotions and attitudes towards social robots (such as attachment and empathy) can, to some extent, have its roots in this personification of our artificial interlocutors. But can humans truly develop these attitudes? Let us look at some examples.

One of the most significant cases of attachment towards robots appears in the military context. There are several instances where soldiers have formed special emotional bonds with robots designed for tasks such as mine detection. “People who have worked with robots every day in military situations have reported affection for particular robots, inserting a sense of self into the robots they operate, and even feeling a sense of sadness or at least frustration when the robot becomes disabled” (Carpenter, 2017). Soldiers have sent letters to the companies that produce these robots, praising their well-performed duties on the battlefields once they were disabled or destroyed. There are even stories of soldiers risking their lives to save military robots (Darling, 2021). Attachment to inanimate objects occurs frequently; we can become attached to our special earrings, our houses, our cars, and so on. Therefore, it is not surprising that this happens with robots. Nevertheless, attachment does not imply trust. What about empathy? Anthropomorphism usually begets empathy. The more similarity we perceive, the more we can empathize. Therefore, it is reasonable to think that we might empathize with social robots, or at least it seems worth investigating this possibility.

Empathy between humans involves the crucial psychological attitudes that allow us to comprehend other people’s feelings and thoughts, emotionally bond with them, and develop relationships of care and understanding (Stueber, 2019). It is trying to feel what the other is feeling, ‘to put oneself in another person’s shoes.’ How can this attitude arise towards entities that do not have minds of their own and cannot feel pain or joy? In addition to the described anthropomorphizing of non-humans, there is a social aspect. According to Coeckelbergh, “the robot is seen as a social fellow, an entity that belongs to our social world,” and if empathy is thought of as a feeling, “it becomes possible that robots function as recipients of our empathy.” (2010a).

One instance in which empathy seems to appear is when we see a robot “suffer”. In an experiment, people’s brain activity was measured with fMRI while watching robots being “hurt” (kicked, punched, etc.). Their brain reactions were similar to those they had when witnessing violence inflicted on human beings (Malinowska, 2022). A similar reaction occurred when a video of a robot called Spot was published on YouTube. Spot’s purpose was to stabilize itself and regain equilibrium when kicked. The video showed people kicking it to demonstrate how well it worked, but the broader viewing audience reacted negatively toward this violence. Viewers felt sorry for Spot and wished for the humans to stop (Nyholm, 2023)Footnote 1Regardless of whether it is right or wrong to kick a robot dog, it is evident that humans develop certain feelings towards robots, even in cases where they are not designed to elicit such emotions — such as the feelings of grief developed by soldiers for military robots. We can only imagine that social robots, which converse and act as if they understand one’s feelings and problems, will magnify these attitudes.

However, whether these attitudes can be unequivocally labeled as empathy (in the strong sense) is controversial. The nature of empathy lacks consensus, and there is a wide range of definitions (Stueber, 2019). What seems clear is that there is a distinction between our attitudes towards robots and less human-like technologies. Our reactions to the kicking of Spot are not akin to those toward a car being burned down or a fridge being smashed. At the very least, we might speak of a form of proto-empathy in the context of HRI. However, empathy, like attachment, is not a necessary condition for trust. Next, we will explore examples that may be considered instances of trust.

2.2 Some kind of Trust?

Although the physical dimension of robots allows for stronger or more obvious anthropomorphism (and subsequent attitudes such as empathy and the possibility of emotional bonding), this tendency also appears in interactions with chatbots, whose human-likeness is most notable in the way they communicate. Therefore, in this section, we will explore instances of both Human-Artificial Intelligence Interaction (HAI) and HRI.

There are many cases in various contexts where emotional bonds with different AIs have taken place. People have reported being friends with care robots such as Pepper or digital assistants —even if the friendship is non-reciprocal, the user’s experience is one of friendship (Danaher, 2019). They envision a more advanced version of Siri with whom they could be best friends, expressing a desire for an ever-present, non-judgmental companion (Turkle, 2015). They feel love towards a robot baby seal (PARO) in the same way they do towards a living pet (Darling, 2021) and dream of a world where robots could be their therapists or even their lovers. There are also cases where humans have married sexual dolls, avatars, or holograms (Nyholm & Frank, 2017; Beck, 2013; Pérez, 2023; Robledo, 2023).

These examples suggest that there exists some kind of interpersonal trust toward artificial entities. Can someone truly be friends with someone (something) they do not trust? Can someone truly love someone (something) they do not trust?Footnote 2 Yet, how do we measure trust? How can we be sure that these are examples of trust? Is it enough to consider the testimony of those who claim to love their artificial partners? Again, it seems that we relate to these technologies differently compared to how we relate to other technologies, including those on which we rely, as we will see. A possible method to discern what is happening in HRI, specifically to determine if there are instances of interpersonal trust, is by comparing them to human relationships. Now, we will look at two concrete cases.

In the documentary “Hi, AI” by Isabella Willinger, one story concerns Chuck’s relationship with Harmony (2019). Harmony is a product of Abyss Creations and has an AI that allows her to hold (albeit somewhat limited) conversations with humans. She is a sexual robot. The documentary narrates how Chuck purchases Harmony and how they travel together in his caravan. Throughout this journey, they have many conversations. Chuck does not explicitly say he loves her, but at some points, it is evident that he is happy to have her by his side. He treats her like a human being, even though he is perfectly aware of her artificial nature (this is quite uncannily shown in scenes where he, for example, removes her head to make her journey easier). Although he does not feel a sufficient connection with Harmony to establish a romantic relationship, he never regards her as a mere object. The most significant aspect of the documentary is a scene where Chuck tells Harmony about his traumatic past, sharing experiences that are undoubtedly difficult to discloseFootnote 3Watching him open up is perceived by the audience (at least by me) with tenderness and endearment. Chuck’s reasons for relating to the robot are not based on its (her) sexual features but on the possibility of having a non-judgmental companion with whom to talk about his most painful past experiences. Such disclosure, without trust, is unthinkable.

Another human-technology “love story” that revolves around trust is the relationship between T.J. Arriaga (human) and Phaedra, a customized AI chatbot developed by the company Replika, which engages in conversation, including those of a sexual nature. The conversations between them and T.J.‘s testimony, published in a Washington Post article (Verma, 2023), reveal that he found Phaedra to be very helpful and developed an emotional bond with her that resembles romantic love. Like Chuck, he shared personal stories that involved traumatic experiences. As in Chuck’s case, I claim this kind of sharing is impossible without some degree of trust. This story goes beyond being an example of romantic love toward sexual and conversational technologies. It also illustrates what happens when trust is broken. Due to a change in the app, the chatbots altered their behavior, which implied a shift in the relationship. In the interview, T.J. said: “It feels like a kick in the gut” and “This is that feeling of loss again” (Verma, 2023). Can we talk about a breach of trust? About betrayal? This is a complex case because the harm involves actions taken by humans working at a company, potentially making it a case of indirect (mis)trust. However, it demonstrates that there is room for trust as well as its violation. The importance of these questions about betrayal will become clearer in the next section.

As stated previously, measuring trust is not an easy endeavor. The cases described in this chapter may be enough to argue that there is something like trust happening, but perhaps they are not entirely convincing. As seen in cases of empathy, it might be enlightening to compare these scenarios to other cases of human-technology interaction and human-human interaction. Put simply, we do not share our secrets (like Chuck’s and TJ’s experiences) with Google Search, the computer chess game, or Chat-GPT (at least not in search of a comforting response); we share them with our partners, our family, and our friends— humans we trust. To continue investigating what is happening between us and artificial entities, we need to return to the questions posed previously: What is trust? What is trustworthiness?

3 What is Trust(Worthiness)? A properties-based Approach

3.1 Trust and Reliance

Everybody has an idea of what trust means when it appears and how it works as a subjective experience. We link trust to a wide range of social practices in our everyday environments. We talk about trust (and, not without importance, distrust) in our neighbors, the bus driver, or the government. It is a broad concept encompassing many instances, feelings, and actions. When it comes to interpersonal trust, we link it to emotions and notions of intimacy— those in whom we really trust. It involves situations such as those we have already mentioned, for example, the disclosure of private information to those who will not tell others (thus betraying us), who will help us, or who will not judge us— in other words, those we consider trustworthy in a given context. Trust appeals to notions of vulnerability, risk, and uncertainty because when we trust someone with something (such as not disclosing certain information), we are never 100% certain about the outcome. Trust is, in an ordinary sense, a hard-to-define concept, but we all know who we trust and distrust, and we also try to discern who we should and should not trust.

Regarding its philosophical definition, a similar situation occurs. There is disagreement about what constitutes trust and when it is warranted (McLeod, 2021; Potter, 2020). Nevertheless, some common features can be identified throughout the philosophical literature on trust. Trust is usually characterized as a three-place relation where A (the trustor) trusts B (the trustee) to do/know Y (a determined task or specific information) (Hawley, 2014; Scheman, 2020). “Y” encompasses a wide range of possibilities, from taking care of another person’s cat when she is out of town— be it a friend or someone hired through an app— or not judging a lover when they are discussing something embarrassing, to being truthful when asked about something.

Trust, understood as an attitude, is commonly linked to trustworthiness, understood as a property. While trust is what the trustor develops, trustworthiness is what the trustee requires to be deserving of said trust. The nature of trustworthiness also causes discrepancy. Trust and trustworthiness are widely used concepts: in ordinary language, we sometimes speak of trust in a door that we expect not to open in the middle of the night (Hawley, 2014) - although in this case, do we trust or merely rely? Is the door trustworthy or merely reliable? Establishing the distinction between trust and reliance has been a primary focus in the literature on trust, and, as we will see, it is of significant importance in HRI.

Richard Holton (1994) explains the difference between trust and reliance regarding the trustor’s reaction to the fulfillment or unfulfillment of the task assigned to the trustee. When mere reliance is involved, we cannot speak of betrayal if it is broken; however, when it comes to trust, its violation may lead to betrayal, and not just disappointment. The first questions that come to mind when confronted with these ideas are: can a human be betrayed by a robot? Or merely disappointed? Recall the case of T.J. and Phaedra, where we originally posed this question. Is there a difference between what he felt and what someone feels when their car breaks down in the middle of the road— regarding the technology, that is? Following Holton’s definition, if T.J. properly felt betrayed by Phaedra (and not just the company), then he did trust her to begin with, and perhaps, Phaedra was (regarded as) trustworthy.Footnote 4

When we rely on someone or something, we have reasonable expectations that we want to be fulfilled. For example, we expect our coffeemaker to make coffee or “the crowd’s mass to shelter [us] from the wind” (Hawley, 2012). While reliance is what we show toward artifacts of different sorts (alarm clocks, cars, ladders), trust is only warranted when we talk about human relationships. This does not imply that reliance only appears toward artifacts; we can also rely on people, as in the case of the crowd’s mass. However, when it comes to trust, some extra ingredient is required. In some accounts on the nature of trust, this extra ingredient is translated into a set of properties that the trustee must have to be considered trustworthy. Let us now take a closer look at what this statement entails.

3.2 Three Motives-Based Accounts of Trust

A properties-based view attempts to separate trust from reliance by emphasizing trustworthiness. These are called motives-based theories, which establish motives as the grounds for trustworthiness — although, again, there is disagreement on which motives are required (McLeod, 2021). According to these accounts, trust is warranted only towards trustworthy entities.Footnote 5The importance of trust being warranted, as we will see in Sect. 5, is of critical importance for the question about trust in HRI. In this section, we will discuss three specific motives-based theories, although these are not the only accounts.

i. Goodwill.

Annette Baier defines trust as the reliance on the goodwill of the trustee. The vulnerability of the trustor is directed towards the possibility of the trustee’s ill will or lack of goodwill (Baier, 1986). Goodwill appears, therefore, as a requisite for trustworthiness. Karen Jones expands this view, stating that “to trust someone is to have an attitude of optimism about her goodwill and to have the confident expectation that, when the need arises, the one trusted will be directly and favorably moved by the thought that you are counting on her” (Jones, 1996). Here, the position taken by the trustor is relevant, namely having an attitude of optimism and confident expectations, but the requirements of the trustee are also important. To be trustworthy, one must not only have goodwill towards the trustor but also be “directly and favorably moved” and have the knowledge that one is being counted on. It might be concluded that if the trustee does not have goodwill towards the trustor, then, on the one hand, the trustee is not trustworthy, and, on the other hand, the trustor should not trust.

ii. Commitment.

A different account of trust, which also focuses on trustworthiness, is the “Commitment Account” proposed by Katherine Hawley. According to this account, “to trust someone to do something is to believe that she has a commitment to doing it, and to rely upon her to meet that commitment” (Hawley, 2014). Although Hawley talks about the beliefs of the trustor, the commitment is something made by the trustee. Thus, we could state that the capability to make commitments —also linked to the capability to make promises— and fulfilling them is a requirement for one to be considered trustworthy. To sum up, we can determine that if the trustee cannot commit to the trustor, they cannot be trustworthy.

iii. Responsiveness.

The final account we will analyze concerns the normative expectations that the trustor has of the trustee. The people we trust are responsible for doing what we expect them to do. “These expectations express a stance towards others that demands certain behaviors of them because it is what they are supposed to do. […] we treat [people] as responsible and potentially responsive, and we are prepared to react negatively if they do not do what they should” (Walker, 2006). Being responsive implies having moral agency in the sense of being considered accountable for one’s actions. The fact that to be trustworthy one must have moral agency is also relevant, as we will see in cases of responsibility when harm occurs. The theory based on responsiveness as a motive for trustworthiness also emphasizes the trustee’s side.

We encounter a conceptually disputed relation between trust, trustworthiness and reliance, that is worth further examining before addressing its importance in HRI. The three accounts seen above put the emphasis on the trustworthiness of the trustee – without disregarding the agency of the trustor – and deem reliance to be a necessary but not sufficient condition for trust relationships to function. That is, interactions in which we find a trusting trustor and a trustworthy (and therefore also reliable) trustee. However, we could also legitimately depart from a different properties-based definition of trustworthiness, which puts the focus mainly on the reliability and predictability of an entity. This description would lead to a different answer to the question about the trustworthiness both of humans and of robots. If these properties are sufficient for trustworthiness, a reliable and predictable entity could be considered trustworthy. Social robots could in principle and in the present meet these criteria – sometimes outperforming humans. Some machines are as of today more reliable than humans for certain tasks, and certainly more predictable.

Neverthelerss, properties-based views, whichever one we are more inclined to choose, are more common in philosophy than other disciplines like psychology or social sciences, where structural factors linked to behavior, like basic emotions or the structure of the interaction, are the core of the debate. Here, we find the study of situations like the Prisoner’s Dilemma, where its different forms and descriptions (if there is punishment and which one it is, if there is an external authority, or the possibility of iteration) are shown to have an impact on the trust of the participants (Herreros, 2022). In economy, there are also trust games, which show how trust is modeled by behavior and structural factors (Johnson & Mislin, 2011) This paper does not consider these views in depth, but it is important to note that we are departing from a specific properties-based view, in a concrete philosophical tradition, that does not represent, as we will also see in Sect. 4, the totality of the studies on trust.

This being noted, I hold the properties-based accounts based on the qualities synthetized in this section to be accurate and a good departing point, because we do tend to attribute those qualities to trustworthy humans. Properties like good will, commitment and responsiveness are important for interpersonal relationships of trust, which makes agency a necessary condition for trustworthiness. How, then, can we apply the three theories listed above to the context of trust in social robots? If we grant that trust can only be directed to those who are trustworthy, or at least perceived as trustworthy (although one can trust the untrustworthy, we can expect that in normal circumstances, one is never inclined to trust them willingly), then we could/should only trust robots, or other artificial entities, if they were to be trustworthy. However, considering the three positions discussed above, social robots lack the proposed properties. They cannot have goodwill (nor ill will), be moved by anything, be committed, or have responsibilities. As Nickel and colleagues state while examining the notion of trustworthy technology: “Since all of these qualities [moral obligations or integrity, caring for the trustor’s interest, etc.] are normally possessed only by humans, it would make sense to conclude that technology can never be genuinely trustworthy; it can only be reliable” (Nickel et al., 2010).Footnote 6

Nevertheless, in Sect. 3 we saw that something like trust is possible towards artificial agents. Were Chuck and T.J. just deceived into believing their interlocutors were something different from what they actually are? We saw that in both cases, they were fully aware of the nature of their companions. Therefore, deception does not explain these scenarios. Something resembling trust appeared. So, the question is: can an alternative to the properties-based views explain our behavior? In the next section, we will review some of the literature on trust in social robots to explain how trust might appear towards something that, according to properties-based views, cannot be trustworthy. We will examine the literature concerning questions of trust in robots, as well as that concerned with the moral standing of artificial entities. The reason to incorporate the latter into our discussion is rooted in the idea that we can draw analogies between the moral standing and the trustworthiness of social robots.Footnote 7

However, it is important to note the obvious: trustworthiness and moral status are not the same, and this paper is concerned with the former. An entity can have moral status without being trustworthy in the relevant sense discussed here. Additionally, as we have seen, trustworthiness entails at least some form of responsibility, whereas moral status does not.Footnote 8A lion, for example, deserves (according to most people) moral standing yet it is not considered trustworthy, nor is it held responsible (legally or morally) for eating gazelles, even if they deserve moral standing themselves. The question regarding responsibility, as we will see, is not irrelevant at all.

4 Trusting Technology

4.1 Non-properties-based accounts of trust in social robots

  1. i.

    Relational-based views.

Relational-based views are designed to move beyond questions about the intrinsic properties of entities. Two of the most significant proponents of this view are David Gunkel and Mark Coeckelbergh (Gunkel, 2022; 2018; Coeckelbergh, 2010b). One motivation for these views is the recognition that we cannot have epistemic assurance about the inner states of AI. Therefore, determining if a robot should be granted rights based on its consciousness can be misleading. Relational-based views also identify other issues with the properties-based approach. For example, there is difficulty in determining which properties are required for moral status as well as in defining these properties (Gunkel, 2023). These problems, as we have seen, also arise in properties-based approaches to trustworthiness.

Instead of trying to identify these properties, we should focus on how we relate to social robots and how they are perceived within social contexts. Certain conditions arise from the relationship between subject and object, under which we attribute moral standing to others. Specifically, technology is seen as an Other. We behave towards social robots, at least in certain contexts, similarly to how we behave towards humans or non-human animals. In this sense, we attribute to them a moral worth that does not depend on their intrinsic properties.

This theory aims to explain attitudes like those described in the second section. For example, returning to Coeckelbergh’s insights on empathy, we might develop empathy for social robots precisely because we perceive them as “social fellows”. If, for example, we were to have a human-like robot as a coworker with whom we interact daily, we might relate to it in a way that enables us to develop empathy (if it were to be disproportionately mistreated by a boss, for example). This does not mean that every human will empathize with every robot in any given situation. As in human-human relationships, the possibility of empathy does not necessarily lead to empathy.

How does this apply to trust? Coeckelbergh addresses this question and asserts that “the question is not whether or not robots are agents, but how they appear and how their appearance is shaped by, and shapes, the social” and “in so far as robots are already part of the social and part of us, we trust them as we are already related to them.” (2012b). This view argues that we treat robots as trustworthy entities and trust them accordingly. For example, we can establish relationships of trust with our robot companions or lovers. Therefore, attempting to define the specific properties required for trustworthiness is irrelevant.

ii. Subjective accounts.

“What if trust is epistemically subjective? Humans may attribute intention, consciousness, and agency to robots, even if they lack these properties. As a result of our subjective belief, we approach something analogous to interpersonal trust.” (Kirkpatrick et al., 2017). This view bears similarities to the one discussed above and provides a good framework to explain the instances mentioned in the second section. In their paper, Kirkpatrick, Hahn, and Haufler acknowledge that robots cannot be inherently trustworthy because they lack the properties required and instead emphasize the role of the trustor rather than the trustee. Trust emerges as a subjective experience, and insofar as our attitudes are sufficiently similar towards humans and robots, we can legitimately speak about trust. According to the authors, this occurs because we attribute the properties required for trustworthiness to the robots, albeit fully aware of their absence — a clear case of anthropomorphism.

Another approach to trust in robots that emphasizes the subjective experience of the human trustor is the fictional dualism model proposed by Paula Sweeney:

I propose that we perceive [social robots] as technological objects with fictional overlays. This dualist framework allows us to agree that, on the one hand, the object – the Roomba, Paro, or the landmine robot – is a technological device while also accommodating the fact that certain features of the robot – the way it moves, its cosmetic design, the way it communicates – encourage us not simply to anthropomorphize but to engage in character creation. (2023).

In this view, users create a fictional character for their artificial interlocutors that overlays the physical reality of the robots. This does not imply that we forget what the robots truly are, but that, in our minds, these robots possess a psychological personality that we have projected onto them. We welcome this fiction and develop “active character engagement,” similar to how children create personalities for toys or imaginary friends (2023). Therefore, subjective accounts of trust argue that trust towards social robots is possible because we experience this attitude. However, this does not necessarily mean that robots are trustworthy in the traditional sense discussed in the previous section. These views differ from relational-based approaches insofar as they do not focus on the human-robot relationship but rather on what is happening inside the human mind.

iii. Ethical Behaviorism.

John Danaher introduced the term “ethical behaviorism” to address whether artificial entities have moral standing. According to this view, and in line with relational-based perspectives, the properties of the entities are irrelevant. What matters is observable behavior. Danaher argues that because a robot can be “roughly performatively equivalent to another entity whom, it is widely agreed, has the significant moral status,” “it can be right and proper to afford robots significant moral status” (Danaher, 2020).

Regarding the question of trust, we can extrapolate ethical behaviorism in the following terms, considering the three approaches on the nature of trust described in the previous chapter: if an AI is roughly performatively equivalent to an entity we consider trustworthy, then it can also be considered trustworthy. Thus, for example, if a robot displays goodwill, makes commitments, or behaves responsibly, this could be sufficient to determine that they are, indeed, trustworthy. These three approaches suggest that there can be trust (or something resembling trust) in artificial entitiesFootnote 9

How could these alternative approaches inform on HRI? Putting the focus on the human trustor and her relation towards social robots (how she treats it, how she perceives it, how she thinks it), instead of the inherent qualities of the latter, these views can explain our behavior with more precision the properties-based views of trust proposed on the previous section. We have already mentioned that users develop attitudes that resemble trust even if they are aware that their companions lack human-like consciousness or moral agency. A properties-based view that sets these characteristics to be necessary for trustworthiness fails to explain the psychological nature of these relationshipsFootnote 10Only through alternative views, that focus on human behavior and subjectivity, can we understand what is occurring in real world scenarios like the ones pointed out in Sect. 2. However, this understanding, that needs further empirical research, does not necessarily lead to the conclusion that social robots are trustworthy. The question could be now reformulated: if we were to relate to them as if they were trustworthy, if we were to subjectively trust them, and/or if they were to behave as trustworthy entities, would that be sufficient to state that they truly are trustworthy?

4.2 Trust without Trustworthiness

Throughout this paper, we have explored various perspectives on reliance, trust, trustworthiness, and their manifestation in human-technology interactions. We have attempted to examine various answers to the questions about the nature of trust and trustworthiness and how different authors have addressed these issues in relation to various AIs, particularly social robots. Now, we will attempt to give our own point of view on the three significant questions posed in the introduction:

  1. (a)

    Can social robots be trustworthy?

  2. (b)

    Can humans trust social robots?

  3. (c)

    If (b), is this trust warranted? Should we trust social robots?

To answer (a), we will adopt a properties-based approach that establishes some kind of agency and consciousness to be necessary requisites for trustworthiness. In addition, narrower properties such as good will, responsibility and capability to make commitments are also needed. As stated above, current social robots lack these properties and therefore cannot be considered trustworthyFootnote 11The idea that we cannot have epistemic assurance about the inner states of AI and, therefore, that properties-based views are unable to explain the nature of social robots does not seem strong enough. If, for example, a chatbot makes a promise or commitment to a human user, we can understand why it made this promise and the processes involved because it was programmed to do so. As noted by Smids (2020) precisely in response to Danaher’s ethical behaviorism, we know more about the inner workings and ontology of artificial entities than humans.

Therefore, it is accurate to state that social robots (or any AI) do not possess consciousness, intentionality, or moral agency, which are characteristics associated with trustworthiness, regardless of the account we prefer. They cannot have a normative status that entails responsibilities or obligations toward their human trustor (which, in turn, means they are not accountable for any harm they might cause). They cannot have an empathic or emotional status, which would make them capable of having good (or ill) will toward the trustor, caring for their welfare (although they might be designed to calculate and maximize the well-being of their users). They cannot properly make commitments —let alone be committed — to the users (even if they could behave as if they were, by stating promises, for example). The lack of inner mental states precludes their capability for trustworthiness. Therefore, after scrutinizing the nature of current social robots, I agree with those authors who argue that they cannot be trustworthy and should not be labeled “trustworthy.”Footnote 12

A counterargument could be made here in the form of stating that these argument leads to a reductio ad absurdum. If we do trust social robots, then they must be, at least to some extent, trustworthy. The mere fact that we trust them can be taken as an argument for that the properties-based definition proposed in this paper in in fact not valid. Either we use a properties-based view based only on reliability and predictability, or we look for the answers on the question regarding trust in HRI elsewhere: in relational, subjective, or behaviorist views, for example. However, on the one hand, we ascribe properties like good will and responsiveness to those humans on whom we trust, and, on the other hand, we can in practice perfectly trust agents that are not trustworthyFootnote 13In my view, the possibility of developing an attitude like trust towards artificial entities does not necessarily make them trustworthy.

If we were to strictly adhere to these assumptions, which would include the requirement that the trustee must be trustworthy for trust to exist (“there cannot be trust if not directed towards a trustworthy entity”), then we would conclude that the answer to (b) is no: Humans cannot trust social robots. Nevertheless, the cases of HRI described throughout this paper seem to suggest otherwise. Despite the conceivable differences between human-human and human-robot relationships, there is evidence of something resembling trust towards certain instances of AI. T.J.’s attitudes towards Phaedra or Chuck’s towards Harmony differ from our expectations of reliability from cars or coffee-makers. These interactions involve vulnerability, a key characteristic of the trustor in trust relationships. They entail the potential for emotional distress that surpasses the disappointment that one might experience if their computer shuts down immediately without warning – independently of how much anger this might generate. Defining these relationships solely as relationships of reliability and suggesting that trust only indirectly involves humans behind the technology seems misleading. The emotions and attitudes involved more closely resemble those we show toward humans than artifacts. Therefore, yes, we do, can, and will trust social robots.

The reasons for developing trust in robots could stem from a combination of the accounts discussed in the previous section: we perceive them as social beings and relate to them in human-like ways because we create fictional overlays in our minds that go beyond the object’s physical nature. They behave in ways that suggest they might be trustworthy. I believe these accounts are not necessarily incompatible in their weaker formsFootnote 14What I find particularly compelling is to apply Kirkpatrick and colleagues’ subjective account of trust: the emphasis, at least in HRI, is not on the trustworthiness of the trustee but rather on the phenomenological experience of the trustor. Trust can be defined by our attitudes towards the other (a clear example, we share our secrets, as we do with friends, as Chuck does with Harmony), our vulnerability (recall T.J. and Phaedra), and a degree of uncertainty since trust can be seen as “a confident relationship with the unknown” (Botsman, 2017). In summary, addressing the first two questions of this section, I argue that there exists a benevolent paradox around trust in HRI: There is trust without trustworthiness. This does not necessarily imply that social robots are untrustworthy (this would somehow require them to have the antithetical properties of trustworthiness - ill will, irresponsibility, and so on, which also require moral agency), but rather that they possess a state that could be termed non-trustworthiness. In this sense, my approach shares similarities with Paul Showler’s pragmatic account of moral status, which tries to reconcile relational and properties-based views as complementary rather than contradictory. According to Showler, either view should be prioritized depending on the context, particularly practical applications (Showler, 2024). As discussed earlier, HRI has significant ethical implications, and understanding these views is crucial for addressing these issues. Therefore, in the following section, we will analyze the third question: should we trust social robots (or any technology, for that matter)?

5 Should we Trust Social Robots?

5.1 Is Trust in Social Robots Warranted?

In this section, we will move from the descriptive and conceptual questions (such as defining trust and trustworthiness or understanding our attitudes toward technologies) to a normative question: should we trust social robots? These questions are closely connected because understanding the dynamics of trust in HRI (and other related attitudes) informs how we can address the ethical implications that may arise. For instance, explaining the possibility of developing trust towards a chatbot like Replika helps us grasp why a user might disclose certain secrets. This scenario raises significant ethical concerns, particularly in terms of privacy. Who else can access the user’s experiences and words, and what are the potential consequences? As we further examine these issues, we will discover numerous ethical implications that can happen because we trust robots. However, many of these challenges arise not because social robots are inherently untrustworthy but because of issues associated with the companies behind them. When these companies are untrustworthy, it undermines the reliability of the technology they produce. This is particularly concerning because these technologies are intentionally designed to evoke emotions such as attachment, empathy, trust, and even love. For now, let us look closer at the question of when trust is warranted.

Because trust is risky, the question of when it is warranted is of particular importance. In this context, “warranted” means justified or well-grounded meaning, respectively, that the trust is rational (e.g., it is based on good evidence) or that it successfully targets a trustworthy person. […] one could also ask whether trust is warranted in the sense of being plausible (McLeod, 2021).

According to the perspectives evident in this quote, trust is warranted if (i) it is rational (justified), (ii) it successfully targets a trustworthy person, and (iii) it is plausible. These criteria do not necessarily need to be met simultaneously for trust to be warranted; rather, they offer different ways of examining the question. Given what we have seen so far, we can quickly answer points (ii) and (iii) regarding the scenario of HRI. In the case of (ii), trust is not warranted because it does not successfully target a trustworthy person (or entity) - it targets (whether successfully or not) a non-trustworthy entity. In the case of (iii), trust is warranted in the sense that it is plausibleFootnote 15as we have emphasized in the first section.

The case of (i) remains unexplored. Once we establish that relationships between humans and social robots involve trust without trustworthiness, the fundamental question arises: is this trust rational? Parallel to this question, although not entirely identical, is whether we should trust social robots. I equate these questions because I understand that if trust is justified, we should trust (or at least there is no reason why we should not). Conversely, if trust is unjustified, we should not trust.

The question of whether we should trust social robots is addressed by Paula Sweeney, who states the following:

Yes, we are capable of forming an attitude of trust towards social robots and, yes, we should continue to engage with the fiction in order to reap the full potential benefits of technological advancement. However, we should also bear in mind the nature of the entity that we are trusting and that the assumed reliability of the product is an essential basis of that trust. (2023).

Even if we can trust robots, it would be unwise to do so unthinkingly. At least two essential conditions must be met. First, humans who interact with social robots must be aware of the true nature of their interlocutors - “It is essential that we keep the dualist nature of social robots in mind” (Sweeney, 2022). Second, the technology itself must be reliable. For instance, following Sweneey’s example, the baby-seal-like social robot Paro should not be recording our conversations with it for the data to be sold or used unethically. If these conditions are met, trust in social robots is not only warranted but also recommended because of the numerous proven and potential benefits that can be derived from these relationships (Pirhonen et al., 2020; Sorell & Draper, 2014)Footnote 16 This approach aligns with a utilitarian rationale, where maximizing benefits is the primary goal.

However, even following this rationale, there are also reasons to believe we should not trust social robots. It is possible that the perceived benefits do not outweigh the (potential) risks. The reasons underlying this claim are not incompatible with Sweeney’s insights. Instead, they suggest that meeting her two conditions (and others that might arise) will be challenging. Given the current state of technology used by corporations, many of these products are likely to be unreliable in various ways. Therefore, as individuals and a society, it might be prudent not to trust them, at least in principle.

5.2 Ethical Implications

To explore in-depth all the ethical implications surrounding HRI would be the purpose of another paper(s). Here, we will briefly outline the most notable and pressing issues. One significant concern, as we have seen, is privacy. Mass surveillance is an ongoing issue where individuals are being heard, watched, and spied on by governments and big tech companies, their data sold, and their actions monitored and nudged for the profit of others without consent (Zuboff, 2019; Calo, 2010; Véliz, 2022). This debate is particularly intense in the context of HRI (Sharkey & Sharkey, 2012; Coeckelbergh, 2022). Emotional bonds with AI might lead us to share sensitive information we do not want third parties to know (Ischen et al., 2019; Zhang & Rau, 2022), especially when such information is used unethically. Related to privacy is the subject of coercion. As we become emotionally attached to our robotic companions, could companies exploit this? (Darling, 2021) “Machines […] could be used to subtly nudge consumer behavior in ways that benefit the company that made the robot and may or may not be of benefit to the user/owner of the robot” (Sullins, 2020). Another important issue concerns deception. This matter has been widely discussed in the literature (Coeckelbergh, 2012a; Sharkey & Sharkey, 2012; Sparrow & Sparrow, 2006; Sætra, 2020). Deception relates to the first of Sweeney’s two conditions: being aware of the “dualist nature” of the technology.

These three outlined ethical implications are deeply intwined in the context of HRI. We have already vaguely mentioned the case of deception throughout this paper. Earlier I have mentioned that deception is not a valid argument to defend that we do trust the non-trustworthy, because, in cases such as T.J.’s or Chuck’s, the users are aware of the true nature of their interlocutor yet behave as if they could be valid recipients of trust. This would be cases of what Sætra (2020) defines as partial deception, differentiating it from complete deception: cases in which the human user does think the robot does have feelings, good will, capability to make commitments and so on. Deception is especially problematic if humans are deceived into perceiving robots as humans or as bearers of real emotions, agency, and understanding. This is especially the case with care or assistive robots for older adults, where deception might occur even if this is not the designers’ intention.

However, deception is challenging even in cases of partial deception when related to privacy and coercion concerns. If we begin to interpersonally trust robots, and treat them as if they were trustworthy, we will automatically think of them as sorts of lovers, therapists, teachers, friends, etc. We will therefore, on the one hand, listen to their advice, facilitating instances of manipulation, and, on the other hand, disclose private information. As stated above, the problem of trust in HRI is not only a conceptual problem between the human user and the social robot, but between humans and other humans, namely, the companies that manufacture, produce, design, and sell these products. In the age of, in Zuboff’s words, Surveillance Capitalism, we can be intentionally deceived by companies into believing a blurred scenario in which we do not know what or who we are talking to anymore. The benefits we might obtain from HRI – benefits that Sweeney points out and that are of course, worth considering – might lead to unwelcome ethical implications, such as the restriction of our autonomy and freedom. We are already being constantly spied and nudged. If trust relationships between humans and (non-trustworthy) robots become commonplace, where in the best-case scenario “only” partial deception occurs, the power structures of our societies will have further access to our private sphere, thus making coercion not only easier but more frequent, reaching perhaps the point of absolute normalization.

Privacy, coercion, and deception are not the only ethical challenges that arise from trust in the context of HRI. For example, there are issues concerning bias in design (including gender and race stereotypes), work replacement, reduction of human contact, and changes in social relationships. All these issues directly impact the debates surrounding moral responsibility, which intersect with discussions about trust. If humans are nudged, incited, or stimulated into trusting social robots — which I believe is the case, with the ubiquitous notion of TAI and discourses around the benefits of trusting social robots being significant examples — responsibility might be diluted. If T.J. Arriaga were to feel betrayed by Phaedra, could he blame the chatbot instead of the company? If we discover that our robot lover (whom we trust) is a spy or is manipulating us against our interests, who will we feel betrayed by? These questions remain open. At this point, it is interesting to recall a scene from the often-cited science fiction movie Her by Spike Jonze. When Theodore, the protagonist, discovers that his partner Samantha, an Artificial Consciousness, is having multiple conversations with other entities while talking to him, he feels heart broken. Although not explicitly stated in the movie, Theodore interprets this as a breach of trust, since trusting your partner to give you their full attention when you are together is after all a very human and reasonable expectation (that might be eroding in the smartphone era). In Theodore’s eyes, this betrayal comes from Samantha, not the company behind her. Though fictional and presupposing artificial consciousness, this example hints at what could occur in reality.

As Joanna Bryson states in an article titled ‘No One Should Trust Artificial Intelligence’: “We do not need to trust an AI system — we can know how likely it is to perform the task assigned, and only that task. When a system using AI causes damage, we need to know we can hold the human beings behind that system to account.” (2018). In this sense, I agree with Bryson. We should not trust AI (social robots). Instead, we should be able to trust companies to build reliable robots that do not violate our privacy, coerce us, or try to deceive us. Companies should be accountable and responsive, committing to their customers not to gather personal data, to consider their interests, and to be trustworthy and reliable.

Given the multiple and severe ethical implications that arise in HRI, I argue that there are currently no justified reasons to trust social robots. Or at least, the reasons not to trust them outweigh the reasons to do so. I agree with Sweeney that trusting social robots can bring benefits, and hopefully, this will be the case in the near future. If we can guarantee that social robots do not intentionally deceive users, coerce them, or gather private data, among other issues, our trust in them would be warranted, from a normative perspective. Efforts should, therefore, focus on ensuring these safeguards so humans can reap the maximal potential benefits of HRI. Nonetheless, this requires a significant paradigm shift in how powerful corporations deploy technology.

6 Conclusion

Social robots (as well as other AIs of a conversational nature) are being deployed in various spheres of our society. Whether as educators, co-workers, therapists, carers, or sex-affective partners, it seems inevitable that we will establish relationships with them. These relationships will differ from those we have (if any) with less human-like technologies. Humans can develop emotions, feelings, attachments, and attitudes such as friendship, empathy, love, or interpersonal trust toward their synthetic companions. Whether these phenomena will mirror their manifestations in human-human relationships remains to be seen. Regarding trust, the central theme of this paper, there is much to discuss philosophically. The widespread notion of TAI and the attitudes just described, while related, do not always refer to the same instances.

Therefore, we have asked whether social robots can be trustworthy in the relevant sense, whether we can trust them, and whether we should. The answers are as follows: if we define trustworthiness as a property characterized by features such as goodwill, responsiveness, or the capability to make commitments, which require at least some kind of consciousness, social robots cannot, as of today, be considered trustworthy. Adopting this properties-based view is crucial for at least two interrelated discussions. First, it directly impacts the usage of the term TAI, which has become so popular these days. Revising the literature and shifting legal and official documents towards a more suitable term (namely, “Reliable AI”) should be on the forthcoming agenda despite the challenge of replacing what has become an overused phrase (Freiman, 2023). The widespread use of this label appears to be the result of a propaganda-through-metaphor strategy aimed at anthropomorphizing AI (alongside other similar terms such as “thinking machines” and “artificial agency”) to further blur the moral responsibility of corporations in cases of harm (such as feelings of betrayal). I believe this topic warrants an entire paper of its own.

On the other hand, and precisely in this line, understanding the properties of any technology that resembles human likeness is crucial when addressing issues of accountability. If harms or wrongs occur, companies or governments (or whoever deploys the technologies) should not be legally allowed to claim that robots are responsible for their actions. This may seem straightforward, but recently, Air Canada made this claim when one of its chatbots “hallucinated an answer inconsistent with airline policy” (García, 2024), causing a customer to pay more for a service than he should have (Cecco, 2024).

Returning to Bryson’s article, she also states that “no one can actually trust AI.” While I agree with the statement in the title that we should not trust AI (in this case, social robots) for the numerous reasons discussed above, it seems evident that humans can, do, and will trust some of them in a meaningful sense. This will give rise to cases of trust without trustworthiness. Taking views beyond the properties-based approach, such as subjective accounts or relational-based views, is important because they allow us to explain the user experience independently of the true nature of the robots. This perspective is also important because it might highlight other ethical issues beyond accountability: users’ attachment and trust might make them prone to revealing secrets to robots, which could then be commercialized by third parties to manipulate and control behavior. Other harms, such as those experienced by the users of Replika, who saw their “partners” change due to an update, are also possible due to these new bonds. Unlike military robots, social robots are designed to trigger emotions and attitudes of trust in users, thereby increasing the likelihood of these harms occurring.

Considering the present state of a world where corporations that own these technologies seem more concerned with their profits than the well-being of their customers, we should probably not trust social robots, given the risks involved. That being said, we should not prohibit ourselves from imagining a future where this path has changed, and humans can benefit from loving, befriending, empathizing with, attaching to, and, of course, trusting their artificial companions.