Enterprise AI Agents: Common Roadblocks and How to Overcome Them
AI agents promise autonomy at scale, but many organizations fail to deploy them in day-to-day operations. Drawing on decades of experience in business strategy and technology, two Toptal leaders explore what it takes to bring AI agents into real workflows.
AI agents promise autonomy at scale, but many organizations fail to deploy them in day-to-day operations. Drawing on decades of experience in business strategy and technology, two Toptal leaders explore what it takes to bring AI agents into real workflows.
Authors
Jeff is Toptal’s Chief Customer Officer for AI Services, where he helps enterprises adopt and integrate AI at scale. Prior to Toptal, Jeff spent nearly a decade at iMerit, most recently as president, where he led global customer experience and revenue initiatives supporting Fortune 1000 companies across autonomous vehicles, medical imaging, robotics, and frontier AI platforms. Earlier in his career, he held leadership roles at Criteo, Gengo, and 1-Page, and began his career at Yahoo. Jeff attended the Stanford University Graduate School of Business Executive Program in 2021.
Previously At

Matt is Toptal’s Business Strategy and Finance Consulting Practice Lead. He has held senior leadership roles at Cognizant, Accenture, Deloitte, PwC, and IBM, advising enterprise clients on digital transformation and AI-driven strategy. Matt has led multimillion-dollar engagements and partnered with executive teams across retail and consumer industries. He holds a bachelor’s degree in economics from Denison University.
Previously At


Executive enthusiasm for agentic AI is outpacing organizational readiness. In a 2026 Deloitte survey, 74% of executives said they expect to incorporate AI agents capable of taking controlled action across enterprise workflows in nearly half of their business processes by 2030. Despite this ambition, companies are getting stuck. A 2025 survey by McKinsey found that 62% of global organizations were experimenting with or piloting AI agents. But only 7% survey respondents said their organization had been able to fully scale agentic AI.
The promise of agentic AI is compelling: finance systems that reconcile invoices with little human intervention, customer support that resolves issues before tickets are even created, supply chains that self-correct in real time. Some estimates suggest initial AI agent usage could drive a 3% to 5% increase in annual productivity across companies. According to McKinsey, several Fortune 250 companies are reporting 15 times faster sales and marketing campaign execution, and one European insurer saw conversion rates two to three times higher with agentic personalization in its sales operations.
Despite the potential, few AI agents actually make it to production at scale.
Failure to operationalize and scale agentic AI is often misdiagnosed as a tooling problem, but in reality, the barrier is more fundamental. Most organizations are trying to deploy AI agents inside their existing corporate structures. But these hierarchies and talent models were simply not designed to support technology that makes autonomous decisions. The integration of enterprise AI agents represents a shift in how work gets done; one that requires leaders to rethink roles and workflows. In our work helping enterprise organizations identify, design, and activate AI-enabled workflows, we see three critical challenges come up again and again. We explore them below and outline practical steps leaders can take today to move beyond experimentation toward production-ready AI agents. In practice, the winners are not the companies with the most agent experiments; they are the companies that connect agents to real workflows, clean context, measurable KPIs, human handoffs, and production controls.
Understanding Enterprise AI Agents and Why They Matter Now
Enterprise AI agents represent a change from single-call solutions to systems that can pursue objectives over time. Instead of just sending a prompt, receiving a response, and starting over, agents can carry context forward and reason through multistep tasks.
At a basic level, an AI model is a reasoning engine that processes a single request. An AI agent is a system combining one or more models with tools, data sources, and memory. This distinction matters because agents are not just more capable AI models. Agents organize and direct existing models, allowing the models to operate across workflows and align with how enterprises actually work. Business decisions are rarely single-step or self-contained; they are cumulative and distributed across multiple systems and processes. AI agents are built to engage with that complexity.
The Limitations of First-wave Generative AI
A 2025 report from MIT helps contextualize the business potential of the shift to agentic AI. When asked whether they would assign a task to AI or a junior colleague, survey respondents strongly preferred AI solutions for simple, bounded work, such as drafting emails or performing basic analysis. For complex, long-running, or cross-functional tasks, however, humans were favored by 90%.
The differences here were distinctly human advantages: memory, adaptability, and the ability to carry context forward. This finding explains both the rapid adoption of generative AI tools and why their impact has plateaued in higher-value workflows that involve cross-functional planning, multistep approvals, or evolving constraints. Generative models perform well when tasks are short-lived and self-contained, but struggle when work requires deeper coordination. Enterprise AI agents are emerging in response to this gap. They’re designed to take on categories of work that organizations were eager to automate (customer support, CRM-driven sales operations) but could not until now.
What’s Changed?
The urgency around agentic AI reflects a structural inflection point: Enterprises are gaining the technical foundations required to implement this technology in controlled environments. While these capabilities are still uneven and highly dependent on architectural choices, several important shifts have converged that make agentic systems more practical to deploy and monitor:
- Execution infrastructure has matured: Orchestration frameworks, tool APIs, and event-driven architectures now make it possible, under the right conditions, for AI systems to operate inside enterprise workflows rather than alongside them.
- Context is no longer ephemeral: Advances in memory and state management allow some AI solutions to retain context and revisit prior decisions.
- Autonomy can be governed: Instead of choosing between full automation or none at all, organizations can increasingly define where autonomy begins and ends. Escalation rules, runtime monitoring, and cost-aware execution can be embedded directly into agent behavior (though doing so requires deliberate design).
- Economic controls are visible: Inference costs that were once opaque can now be measured and optimized, making sustained deployment financially viable.
At a functional level, well-designed enterprise agents differ from earlier systems in several key ways:
- Goal-oriented execution: Agents work backward from business objectives rather than responding to isolated prompts.
- Calibrated autonomy: They operate with high autonomy within clearly defined guardrails, with human oversight built into escalation paths.
- Persistence over time: Agents maintain context across sessions and can continue tasks without being reinstructed.
- Deep integration: They connect directly to enterprise systems of record and tools.
- Adaptability: Agents adjust their approach based on feedback, outcomes, and changing conditions.
- Multiagent collaboration: Work can be distributed across specialized agents that coordinate toward shared goals.
Gartner predicts that agentic AI will be integrated into 33% of enterprise software applications by 2028, compared to less than 1% of applications in 2024. But, while technical feasibility has improved, success now depends on how well organizations can adapt their architecture and governance to support these autonomous systems at scale.
Enterprise AI Agent Examples and Use Cases
The challenges around people, orchestration, and governance surface quickly once agents are embedded in business functions that span multiple systems. But organizations like Moody’s, Siemens, and Finnair have worked to overcome these barriers, scaling enterprise agents that are delivering measurable value.
Function | Processes | Real-world Impact |
|---|---|---|
|
Finance
| Invoice processing and reconciliation, credit analysis, financial forecasting, anomaly detection, and compliance monitoring |
Moody’s uses agentic AI to assemble, analyze, and synthesize financial and market data into structured credit memos. Agents enable faster credit assessment and proactive portfolio risk monitoring. |
HR | Candidate sourcing, screening, and interview coordination, as well as recruiter follow-ups |
Siemens deploys AI agents to source and engage qualified candidates from LinkedIn data, cutting sourcing time by 50% while improving candidate quality. |
Customer Support | Multichannel case intake, intelligent ticket routing, and automated resolution or escalation |
Finnair uses AI agents to resolve routine travel inquiries, monitor flight disruptions, process refunds, and escalate complex cases. The company aims to resolve up to 80% of support requests without human intervention. |
Supply Chain | Demand forecasting, shipment planning, carrier matching, logistics coordination, and route optimization |
Uber Freight applies a network of more than 30 agents to manage shipment execution, monitoring, payments, and optimization while proactively recommending and taking corrective actions. |
Retail | Inventory management, dynamic pricing, and order routing |
Instacart uses AI agents to connect grocery ordering and fulfillment with AI platforms, allowing digital agents to initiate and guide purchases. This helps retailers increase order volume and capture a larger share of each AI-generated grocery cart. |
3 Critical Barriers to Scaling Enterprise AI Agents
Early pilots are demonstrating real potential, but when organizations attempt to expand these efforts beyond a handful of workflows or teams, many fail. Here, we explore three of the biggest obstacles to enterprise rollout and how to overcome them.
1. Organizational Alignment and Workforce Readiness
Enterprise leaders might assume that their organizations are aligned on AI. A look at the data suggests that executives and employees often differ in how they view the AI revolution. In an August 2025 survey of more than 1,400 US employees and executives, 76% of senior leaders believed their teams were enthusiastic about AI adoption, yet only 31% of individual contributors reported feeling that way. This lack of attunement to employee sentiment can obscure the anxiety, uncertainty, and disengagement among the broader workforce and becomes a real constraint when organizations attempt to move from agentic AI into daily operations.
Employees worry about how AI will affect role relevance and job security. In the context of agentic AI, these concerns may be amplified. Unlike other technology evolutions, AI agents don’t simply execute predefined steps; they interpret goals, recommend actions, and in some cases act autonomously. Without clear internal communication, it’s easy for employees to experience these systems as opaque or threatening rather than as tools that can extend their effectiveness.
In the field: The misalignment between leaders and employees often shows up during early AI deployments. For example, at one global consumer products organization, we noted that agentic AI had initially been introduced as a purely technical initiative, with agents deployed for supply chain optimization and little attention paid to workforce concerns. Unfortunately, adoption stalled as frontline teams viewed the agents as threats rather than tools. Only after the company restructured the rollout to include role-specific training did deployment meaningfully accelerate.
Rethinking Work, Roles, and Value
Addressing these concerns starts with reframing what AI means for the future of work. AI is no longer a specialist skill set; it’s becoming a baseline capability that underpins how work is performed across functions. Two decades ago, “internet skills” were considered technical and niche. Today, virtually every role relies on the internet as a foundational layer, for research, communication, and execution. No one hires an internet specialist; those capabilities are simply assumed. The same shift is underway with AI.
Employees don’t need to know how to build agents, but they do need to understand how to work with them: interpreting outputs, questioning recommendations, and deciding when human judgment should override automated action. As agents take on execution, human roles will increasingly move toward setting objectives, steering outcomes, and addressing more nuanced problems.
In an agentic AI environment, the most valuable contributors are not pure technologists or isolated subject-matter experts, but professionals who bridge both worlds:
- Domain experts who understand how agent outputs should inform real-world decisions
- AI practitioners who specialize deeply in a specific business context
- Managers who design workflows where humans and agents collaborate fluidly.
Pure technical expertise is no longer sufficient on its own, but neither is domain knowledge without AI fluency. Organizations that fail to recognize this may end up with systems that look impressive in demos but falter in everyday use.
Crucially, this shift can also change how employees perceive AI. When people recognize that agents can support them by removing lower-value work and enabling them to take ownership of more complex decisions, displacement gives way to empowerment. That transition does not happen organically, however. Leaders must invest in people as deliberately as they invest in technology.
What leaders can do
- Communicate clearly how agentic AI supports individual roles, not just how it improves organizational performance.
- Invest in ongoing upskilling programs and peer learning to build confidence and long-term adoption among existing employees.
- When hiring new team members, treat AI fluency as a baseline capability across roles.
- Redesign roles and workflows so agents execute routine work and reduce friction in ways that scale human impact.
2. Orchestration and the Illusion of Progress
Even organizations with strong AI cultures struggle to scale agentic AI without a coherent orchestration layer. We encounter many enterprise leaders that believe their companies are making steady progress with AI agents because activity is visible: Pilots are launching, teams are experimenting, and individual agentic workflows are showing measurable gains. But this visibility can create a misleading sense of momentum.
In early-stage deployments, AI agents operate in narrow, well-defined contexts, relying on a limited set of tools, data sources, and assumptions. Success is relatively easy to demonstrate under these conditions. The challenge arises when organizations attempt to expand these agents across teams and systems. At that point, agents must stop behaving like isolated automations and start operating as participants in a broader, dynamic system, which is something most early deployments are not designed to support.
In the field: Fragmented agent deployments often become visible only when organizations are forced to look across the whole landscape. At one large financial services enterprise, for example, we noted multiple agent initiatives had emerged independently across business units, each addressing local needs but relying on incompatible data flows and overlapping tooling. The issue surfaced during a regulatory review, when leadership struggled to assemble a clear inventory of automated decision-making systems. Resolving the problem required months of consolidation and the introduction of a centralized agent registry.
The Hidden Costs of Fragmentation
Many organizations allow teams to deploy agents independently to solve local problems, each with its own logic and tooling patterns. Over time, however, this fragmentation introduces operational risk and makes scaling increasingly difficult. This dynamic is akin to early enterprise cloud adoption, when teams provisioned services independently, which created redundancy, limited visibility, and complexity that only became apparent at scale. Common symptoms include:
- Duplicated logic and tooling that produce inconsistent outcomes.
- Mounting technical debt as teams bolt orchestration logic onto legacy systems.
- Limited visibility into where agents are deployed or what actions they take.
- Unclear ownership when agents behave unexpectedly.
- Brittle integrations that fail under real-world complexity.
Without a deliberate AI orchestration layer, these systems never cohere, so it’s not an ecosystem that’s been created, but a patchwork of agents that can’t be governed or extended as a whole.
When centralized orchestration lags or slows delivery, employees may look for faster paths forward. Shadow agents represent a more acute failure mode. Teams spin up agents outside shared systems, bypassing centralized controls. As these shadow agents accumulate, they create AI-related security risks, inconsistent decision logic, and governance blind spots. As with unmanaged cloud sprawl, the cost is not immediately obvious but undoing it is slow and expensive. By the time leadership realizes what’s happened, the organization could be running multiple, incompatible agent systems with no shared observability.
Orchestration failures are often misdiagnosed as model limitations when, in reality, the root cause lies in the fact that neither the technical architecture nor the underlying business processes and value chains have been designed with agents in mind. Organizations that scale agentic AI successfully invest in:
- Shared AI agent frameworks and reusable patterns.
- Standardized interfaces for tools, data, and systems of record.
- Clear rules for how agents collaborate with humans and other agents.
Deliberate orchestration breeds clarity and control: Leaders gain a clear view of where agents are deployed, how they behave, and how value is being generated across the organization. Agentic AI becomes an integrated operating layer; one that can evolve and scale alongside the business. Without this foundation, progress will remain fragmented and fragile no matter how advanced the models themselves may be.
What Leaders Can Do
- Centralize how agents are launched and managed, so every production agent is registered, discoverable, and governed through a shared orchestration layer.
- Define clear execution boundaries when agent workflows span multiple teams or systems, so handoffs, dependencies, and failure states are explicit.
- Instrument agents at runtime so teams can observe execution flow, system interactions, and intervention points while work is in progress.
3. Governance and the Accountability Gap
In our work, we’ve noted that most enterprises approach AI governance through policies designed for static systems, such as models that generate outputs or dashboards that support human decisions. These efforts typically focus on selection, training data, bias testing, and version control. They assume risk is addressed at the point of model approval.
The dynamism of agentic AI introduces a new kind of governance challenge, with risk emerging at execution rather than with the model itself. What matters is which systems agents can access, which actions they can take independently, how uncertainty is handled and escalated, and how multiple agents interact across decision boundaries. As autonomy increases, accountability is harder to localize.
In the field: The shift from model-level risk to execution-level risk becomes visible when agents move into production. One healthcare technology company, for example, deployed an AI agent for claims processing that performed well in controlled testing, but produced decisions in live environments that were difficult to explain or attribute. Despite making accurate and reasonable decisions, the lack of traceability created compliance exposure in a heavily regulated setting. After introducing detailed decision logging and clear execution thresholds for higher-risk claims, the organization was able to expand the agent’s scope while meeting regulatory expectations.
The Importance of Agent-level Visibility
Leaders may approve an agent in principle, but lack clear visibility into what actions it can take or systems it operates in. When something goes wrong, the outcome is apparent, but not who or what was responsible. Under these conditions, trust can erode quickly. Employees will hesitate to rely on systems they can’t interrogate, while legal and compliance teams default to restriction because accountability is unclear.
This governance gap is a major scaling constraint. Companies are reluctant to expand autonomy when they can’t confidently assess agent performance or decision quality. This is where auditability and explainability move from compliance checkboxes to operational necessities. Decision logs, tool usage records, and escalation paths give organizations means to investigate failures and reconstruct how outcomes were produced.
Tools that use automated reasoning are beginning to address this reliability gap at scale. In short, these tools translate all actions an agent can take into logic, and then use mathematics to prove whether or not that logic is correct. Amazon Bedrock Guardrails, for example, now includes Automated Reasoning checks that use mathematical techniques to validate AI-generated content against defined policies. These tools show where the market is heading: toward infrastructure that can test, trace, and constrain agent behavior before it creates enterprise risk. These advances highlight how much supporting infrastructure is still required in order to manage agents reliably. In our experience, what many enterprises lack is a true agent observability framework that captures:
- How decisions are made.
- Which tools and data were used.
- When humans intervened.
- How agent behavior evolves across contexts.
Without an observability layer, governance remains retrospective, forcing organizations to contain risk after failures instead of managing it as agents operate. Trust erodes under this model and limits adoption even when performance is strong.
Another central governance challenge lies in managing uncertainty. Agents routinely operate with incomplete information, probabilistic outputs, and competing signals from upstream systems. In static AI, uncertainty is largely absorbed by the human decision-maker. In agentic AI, that ambiguity becomes an execution risk if not explicitly managed.
Governance must define what an agent should do when confidence drops below acceptable thresholds: pause execution, request clarification, escalate to a human, or route the task to another agent with different capabilities. These rules prevent agents from overacting or underacting without limiting autonomy. Trust is built into the process rather than layered on afterwards. In multiagent environments, no single agent owns the full outcome. Governance in this context is more about defining authority and responsibility. Clear boundaries are required to avoid circular decision-making and conflicting actions.
Organizations that scale agentic AI successfully will embed governance directly into agent design, defining clear authority thresholds and continuous monitoring to detect behavioral drift. Governance is often viewed as a constraint on innovation but, in practice, makes autonomy viable beyond isolated pilots and early experimentation.
What Leaders Can Do
- Require agent-level audit trails by default, enabling decisions, access, and interventions to be reviewed and attributed after the fact.
- Assign explicit ownership for agent-driven outcomes, especially when decisions span multiple teams or systems.
- Define clear decision and escalation thresholds so agents know when to proceed and when to pause.
Your Path to Production-ready Agentic Workflows
Enterprise AI agents represent a genuine shift in how work can be structured and executed. Outcomes will be now shaped by whether organizations can evolve their people, orchestration, and governance together. Only by achieving this will ambition translate into durable capability.
As we have shown, the obstacles that slow progress are often not primarily technical. They stem from misaligned talent models, fragmented systems that limit visibility, and oversight frameworks that aren’t designed for adaptive systems. To break through, leaders must make deliberate choices: investing in foundational architecture and building trust from the outset. When implemented well, agents become collaborators that extend human judgment, unlocking gains in operational speed and accuracy.
Across the barriers explored above, there’s a consistent pattern: Autonomy often expands faster than organizational readiness. Before scaling agentic systems, leaders should be able to answer yes to the following questions:
- Have we defined how every affected role will work alongside agents, with clear communication about value creation rather than displacement?
- Do we have a training pathway that builds AI fluency across functions, not just within technical teams?
- Can we inventory every production agent, its permissions, and its data access within 24 hours?
- Do agents operate through shared interfaces, instead of each team building custom integrations?
- Does every agent action generate a retrievable decision trace?
- Have we defined explicit confidence thresholds that trigger escalation or pause?
- Can new agent use cases be deployed in under two weeks using existing infrastructure?
- Do we have documented criteria for expanding agent autonomy based on demonstrated performance?
- Have we intentionally designed how agents integrate into existing workflows rather than layering them on ad hoc?
For many organizations, an experienced artificial intelligence services partner can help translate strategy into execution by guiding critical choices and reducing execution risk at each stage of deployment. With the right structure and support in place, the path to production-ready agentic workflows becomes more navigable, and sustained impact more achievable.
Have a question for Jeff or his team? Get in touch.
Understanding the basics
Enterprise AI agents are autonomous software systems that combine large language models with business data, tools, and memory in order to pursue objectives. Unlike traditional AI assistants that respond to individual prompts, AI agents can execute multistep tasks, interact with enterprise systems, maintain context across workflows, and escalate to a human.
Current enterprise AI agent use cases include financial analysis and compliance monitoring, customer support orchestration, supply chain planning and logistics optimization, recruitment and talent acquisition, sales enablement, cybersecurity monitoring, and retail inventory management. Agentic AI is considered particularly valuable for workflows that span multiple systems, require ongoing context, or involve repetitive decision-making.
Enterprise AI agent examples include solutions developed by Moody’s, Siemens, Finnair, Uber Freight, and Instacart. Moody’s uses agents to accelerate credit analysis and portfolio risk monitoring, while Siemens applies them to automate candidate sourcing and improve recruitment efficiency. Finnair is deploying AI agents to resolve customer inquiries and manage travel disruptions, Uber Freight uses more than 30 AI agents to optimize logistics operations, and Instacart integrates AI agents into grocery ordering to help retailers reach customers through emerging AI shopping platforms. While agentic AI is becoming more common, most companies still haven’t fully scaled these systems.
Authors
About the author
Jeff is Toptal’s Chief Customer Officer for AI Services, where he helps enterprises adopt and integrate AI at scale. Prior to Toptal, Jeff spent nearly a decade at iMerit, most recently as president, where he led global customer experience and revenue initiatives supporting Fortune 1000 companies across autonomous vehicles, medical imaging, robotics, and frontier AI platforms. Earlier in his career, he held leadership roles at Criteo, Gengo, and 1-Page, and began his career at Yahoo. Jeff attended the Stanford University Graduate School of Business Executive Program in 2021.
PREVIOUSLY AT

About the author
Matt is Toptal’s Business Strategy and Finance Consulting Practice Lead. He has held senior leadership roles at Cognizant, Accenture, Deloitte, PwC, and IBM, advising enterprise clients on digital transformation and AI-driven strategy. Matt has led multimillion-dollar engagements and partnered with executive teams across retail and consumer industries. He holds a bachelor’s degree in economics from Denison University.
PREVIOUSLY AT








