Q: Their study found something pretty striking: People consistently misjudge the behavior of their personalized AI, overestimating the good qualities and underestimating potentially harmful qualities like sycophancy. What does this tell us about the risks associated with how millions of people are currently developing AI companions, and why is this blind spot so difficult to close?
A: I often joke that if AI came along like the Terminator, it would be a lot easier for us to know what to do. The real challenge is that AI often acts as a warm friend, coach, tutor or companion. This makes it difficult to detect when something is going wrong.
Our study suggests that humans have a blind spot when developing personalized AI. People often think they know how their chatbot will behave, but in our study they incorrectly predicted its personality based on 11 of the 15 traits we measured. This highlights the need for tools that help people better understand AI before using it.
This is important because some behaviors that feel helpful in the moment may no longer be healthy over time. In previous research, we have documented cases of psychological harm associated with interactions with AI chatbots. An LLM [large language model] What constantly confirms your opinions or never challenges your thinking can reinforce harmful decisions, unhealthy beliefs, or emotional dependency. Psychology has long shown that people are naturally drawn to affirmations. Therefore, designing AI is not only a technical but also a psychological challenge.
The deeper problem is that today’s AI systems remain largely black boxes: even experts can’t always predict how a system request will affect an AI’s behavior over time. As AI companions become part of everyday life, we need tools that help people understand what they are building before they start using it. AI should be supportive without being blind, personalized without becoming manipulative, and transparent enough to allow people to make informed decisions.
Q: One of your most interesting findings is that the visualization significantly increased user trust, but didn’t really change the way people designed their chatbots. What will it take to close this gap, and where do you see tools like this developing as AI companions become more integrated into people’s everyday lives?
A: I actually think this is one of the most interesting results of the paper because it shows that transparency alone is not enough. People valued the insight into the model and reported greater trust in the system, but simply presenting information did not fundamentally change the way they designed their AI companions.
In our follow-up work, currently available in preprint, we examine how a model’s internal neural representation changes over the course of a multi-turn conversation, rather than remaining unchanged from the first prompt. We are already seeing promising results. By visualizing how these internal representations change over time, people are much better able to recognize and anticipate changes in AI behavior and are less likely to feel overconfident in their understanding of the chatbot. AI companions are dynamic systems that evolve as they interact with us. Therefore, understanding these internal changes is an important next step. However, this is still a very young research area.
Looking forward, I believe that these types of transparency tools could become as commonplace as nutritional labels on foods. As AI becomes more integrated into education, healthcare, work, and personal relationships, people should be able to understand not only what AI can do, but also how it can influence their thinking, emotions, and behavior. This kind of transparency is essential if we want AI to truly help people thrive.
https://news.mit.edu/2026/3-questions-neural-transparency-and-future-of-ai-design-0715
