The engineering team inside Rabobank's Digital Platform tribe has been running hands-on experiments with AI agents and multi-agent architectures to evaluate production readiness — first prototyping workflows in Azure AI Foundry's Connected Agents, then migrating to the Microsoft Agent Framework for explicit, code-based orchestration.
"In this article, 'we' refers to our engineering team within the Digital Platform tribe, where we experiment with AI agents and multi-agent architectures to build practical, production-oriented solutions. The insights and trade-offs discussed are drawn directly from hands-on experiments and prototypes developed during this work." — Buse Yalcinkaya Demir, Data Scientist/Engineer and GenAI community lead at Rabobank.
The motivation came from limits of single-agent setups: as the team added more tools, the agent began confusing tool selection, behaving brittly under complex prompts and losing accuracy over long context windows. Multi-agent architectures — distributing tasks across specialised agents rather than one generalist — were tested to keep roles clear, reduce errors and stabilise performance, with orchestration patterns (sequential, concurrent, group-chat, hand-off, magnetic) chosen per workflow.
Technically, the first round ran on Azure AI Foundry's Connected Agents, a no-code primary-plus-sub-agents setup with built-in tools (Azure AI Search, Bing grounding, SharePoint, Code Interpreter, Azure Functions, Logic Apps). The team then moved to the Microsoft Agent Framework, an open-source runtime that adds predefined workflow order, modular reusable components, OpenTelemetry observability, a developer UI, human-in-the-loop approval, and MCP / A2A integration for external tool discovery — all running inside Azure AI Foundry with RBAC and safety controls.
Recurring pain points: the supervisor agent sometimes diluted or misinterpreted sub-agent outputs and did not reliably trigger sub-agents in low-code setups; sub-agents occasionally outperformed the supervisor; tool calls could be skipped or misused; and low-code orchestration hid what was actually happening. Connected Agents also supported only one instance per tool, made orchestration implicit and nondeterministic, and offered no guarantee that citations survived across agents.
Lessons on agent design: ask first whether an agent is needed at all (use them only when tasks are dynamic, ambiguous and multi-step); start with one agent and minimal tools before scaling; distribute work across specialised agents; design prompts that define scope, role, task and boundaries, allow "I do not know" responses and follow the single-responsibility principle; choose models based on cost, latency and tool-call reliability rather than always picking the strongest; prefer local memory; run continuous evaluation with messy multilingual inputs to test routing accuracy, tool use and reasoning quality.
Looking ahead, the team frames current agent workflows as "promising but not yet stable and reliable enough for production" — the Microsoft Agent Framework is still in public preview and "not yet recommended for running large-scale production GenAI solutions." Next steps are to keep iterating with guardrails and continuous evaluation while validating that agentic approaches truly outperform standard LLM calls for each candidate workflow.