Why IT Operations Leaders Still Don't Trust AI to Take Action
IT operations teams have largely solved the conversational layer of the service desk. A virtual agent handles routine tickets over chat, a voice agent takes after-hours calls, and in some deployments an avatar greets employees at a help desk kiosk. The responses are accurate, and the deflection rate looks strong in the quarterly review.
Then leadership asks the harder question: could this actually resolve the incident instead of just talking about it? Could a coordinated set of agents execute the fix in the background with exactly as much human oversight as the situation calls for, no more and no less? Most I&O leaders, if they are honest, cannot answer that with confidence yet. The technology is capable of taking the action; what is missing is trust in the reasoning behind it, specifically its ability to reliably judge when a person needs to stay in the loop and when the process can safely run on its own.
That hesitation is not really about the interface, whether the request arrives through a chat window, a voice call, or a multiagent workflow running behind the scenes; it is about whether the AI making the decision can show its work, including the work of judging how much human involvement a given action actually requires. In fact, this is something Gartner has recently explored in their Hype Cycle for AI in IT Operations, 2026.
Why Capability Alone Doesn't Earn Trust
Most GenAI systems, whether wrapped in a conversational interface or coordinating a team of agents, reason the same underlying way: probabilistically, over patterns learned from training data, without a transparent chain from input to decision. That is a reasonable tradeoff for drafting a response. It becomes a serious liability the moment the system is deciding whether to restart a service, apply a patch, or escalate an account privilege in a live production environment, and cannot defend why a human was, or was not, part of that decision.
Gartner's own analysis names this directly as one of the primary obstacles facing GenAI virtual assistants: explainable output is what closes the gap between accuracy and reliability, and providers are increasingly addressing it by combining generative reasoning with rule-based technology rather than relying on the language model alone. Gartner raises a related concern on the multiagent side, noting that coordinating and governing multiple agents is difficult and that multiagent approaches without some form of centralized planning are often unreliable. However, tighter workflow control across agents tends to improve reliability at some cost to flexibility.
What Closes the Trust Gap
Organizations closing this gap are not choosing between generative flexibility and deterministic control. They are combining them. Neurosymbolic reasoning, pairing a language model's pattern recognition and language understanding with rule-based, symbolic logic, gives an AI system an auditable basis for every decision it makes, whether that decision surfaces as a conversational answer or as an executed action inside a multiagent workflow. Gartner's own user recommendations for GenAI virtual assistants point the same direction, advising organizations to look for solutions that pair large language models with rule-based logic rather than depending on the language model on its own, as a way to make outcomes both more dependable and easier to defend to a compliance team.
This discipline matters equally on the conversational side and the multiagent side, arguably more on the multiagent side, since an executed action in production carries a larger blast radius than a poorly worded answer in a chat window. Either way, the requirement is the same: reason over the request, apply policy before acting, verify the outcome, and log every step in a form a compliance team can defend later.
That same reasoning is what determines how much human involvement a given action should have, and the answer is not the same for every action. A change touching a regulated system or carrying real business risk may correctly require a person's sign-off every time. A narrowly scoped, well-understood task may correctly require none. What matters is that the system applies whichever answer the organization's own rules call for, consistently, rather than defaulting to full autonomy or full human review across the board.
Gartner's Recognition and What It Signals
Gartner recognizes both of these categories in the Hype Cycle for AI in IT Operations, 2026, naming Openstream.ai as a Sample Vendor in Multiagent Systems, rated Transformational with 1% to 5% of target enterprises currently adopting it, and in GenAI Virtual Assistants, rated High benefit with 20% to 50% market penetration. Recognition across both categories in the same report reflects what I&O leaders are living through directly: conversational AI and multiagent systems are being asked to clear the same trust bar before either one gets a real operational mandate.
How Openstream.ai Approaches This
Eva, Openstream.ai's enterprise AI platform, delivers both sides of this. Eva powers conversational AI experiences, including AI Virtual Agents, AI Voice Agents, and AI Avatars and Digital Humans, as well as the specialized, semiautonomous multiagent systems that carry out the underlying operational work. Both are built on the same neurosymbolic reasoning foundation, combining generative flexibility with deterministic, rule-based control, so that every agent follows the specific processes, policies, and regulatory requirements an enterprise sets for it.
That foundation is what lets an organization calibrate oversight instead of picking one setting for it. Some processes warrant a person's sign-off every time. Others, once well understood and properly governed, can run with greatly reduced intervention or none at all. Openstream.ai does not build its operational AI to eliminate people from IT operations. Eva is designed to apply exactly the level of human involvement each process calls for, and to make that judgment as explainable as the action itself.
Where to Start
The decision facing most I&O leaders in 2026 is not whether to add a conversational front end or a multiagent workflow. Most organizations will eventually need both. The real decision is whether the reasoning underneath either one can be trusted to apply the right level of human oversight to each action, full sign-off where a process demands it, and reduced or no intervention where the rules and the risk profile allow it, consistently and on the organization's own terms. That trust does not come from a larger model. It comes from an architecture that can show its work at every step, including the work of deciding how much a human needs to be involved, and getting that foundation right now is considerably easier than retrofitting explainability onto a system that was never built to provide it.
We are working through these problems with I&O teams now. If you want to understand what a governed, explainable AI deployment looks like in a live IT operations environment, we would welcome the conversation.
Gartner, Hype Cycle for AI in IT Operations, 2026, Cameron Haight, 10 July 2026, ID G00853454.
GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved. Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner's research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.
