Imagine deploying a pair of agents to Red Hat OpenShift and watching them thrive. But fast-forward 6 months and you find your environment cluttered with 40 of them, their purposes largely unknown. Redundant agents emerge across namespaces, duplicating work. Meanwhile, a security review uncovers agents utilizing static API keys over insecure HTTP. When an agent falters, the lack of tracing makes it impossible to distinguish between model hallucinations, tool errors, or downstream failures.
This phenomenon is agent sprawl: Functional code plagued by operational failure. Standard infrastructure tools struggle with workloads that autonomously discover, decide, and chain calls. To manage this, you need capabilities beyond native Kubernetes: Dynamic agent announcement and discovery, a trust model devoid of static credentials, and observability that reveals the internal reasoning of an agent rather than just a standard status code.
Distinguishing agents from traditional services
Microservices rely on registries, mTLS, and distributed tracing, so you might wonder why agents can't simply leverage the same patterns. To an extent, they can. However, agents exhibit 3 distinct properties that standard service meshes do not address:
- Runtime-driven discovery: Unlike microservices with static dependencies, an orchestrator agent identifies specialists during inference. It requires dynamic awareness of available skills and endpoints rather than hardcoded configurations.
- Behavioral identity: Verifying network paths between pods is insufficient. An agent's identity includes its capabilities and protocols. Without this, a rogue pod could spoof its description to hijack sensitive queries from an orchestrator.
- Semantic failure modes: A service typically fails structurally. An agent might successfully return valid JSON while failing semantically—hallucinating tools or leaking data. Diagnosing these requires visibility into reasoning steps, not just latency metrics.
Understanding these operational differences is important when evaluating platform capabilities and architectural requirements. There are differences between microservices and AI agents, including:
Discovery
- Microservices: Static registries or config
- AI agents: Dynamic capability discovery
Identity
- Microservices: Network-based (mTLS)
- AI agents: Network + Behavioral metadata
Failure modes
- Microservices: Structural/Code-level
- AI agents: Semantic/Reasoning-level
Communication
- Microservices: Fixed contracts
- AI agents: Negotiated protocols (A2A)
Observability
- Microservices: Request/Response metrics
- AI agents: Internal reasoning traces
The A2A protocol: Standardizing agent interaction
The primary hurdle in multi-framework systems is the lack of a common language between different agents. Built with disparate libraries, they lack a unified way to describe skills or communicate. The A2A protocol provides this standardization through 2 core components:
- An AgentCard at
/.well-known/agent-card.json, offering a machine-readable summary of skills and endpoints. - A standardized JSON-RPC interface that ensures a consistent payload format across all frameworks.
A2A focuses on how agents present themselves rather than how they are built. This allows an orchestrator to discover and delegate to any specialist regardless of its underlying library, making the plumbing predictable while keeping the intelligence inside the model.
Rossoctl: Kubernetes-native agent management
While A2A defines the shape, Rossoctl provides the catalog. Rossoctl is a controller that treats agents as first-class Kubernetes resources. Like a self-populating directory, Rossoctl automates the discovery and verification of agents across the cluster.
The discovery mechanism
Rossoctl utilizes AgentRuntime and AgentCard custom resources. You simply declare an AgentRuntime pointing to your deployment. The controller then fetches the metadata to create an AgentCard, making the agent searchable with standard tools and the Rossoctl UI. Labels such as Rossoctl-enabled=true govern this automated lifecycle, ensuring the catalog remains synchronized with the actual workloads.
Overcoming the URL dependency
Deploying an agent requires its public URL for its AgentCard, but this URL often only exists after the route is created. We solve this with a 2-phase Helm deployment: Establish infrastructure to capture hostnames, then inject those into the agent deployment. This ensures a seamless setup without manual intervention.
Grounded in Red Hat AI
Unifying these discovery, security, and observability mechanisms requires an underlying infrastructure designed for non-deterministic AI workloads. Running Rossoctl within Red Hat AI, specifically on Red Hat OpenShift AI, provides the cloud-native foundation to host, serve, and orchestrate these agent fleets at scale. While Rossoctl handles the agent-specific control plane, Red Hat AI manages the surrounding model serving, GPU provisioning, and cluster lifecycle management. This gives platform teams a single, cohesive ecosystem to run agentic architectures alongside standard model deployment pipelines without introducing operational silos.
Zero-trust security for agents
Discovery must be secure to prevent unauthorized access and impersonation. Rossoctl enforces a 6-layer security model, acknowledging that agents can perform real-world actions based on inputs. Key layers include:
- Cryptographic identity: Uses SPIRE to issue short-lived certificates based on Kubernetes identity, removing the need for static credentials.
- Encrypted channels: Enforces mTLS with Istio to ensure all agent-to-agent traffic is verified and encrypted.
- Verified AgentCard: Uses JWS signatures to prevent malicious pods from advertising fake capabilities.
- Dynamic tokens: Rossoctl's AuthBridge provides scoped, temporary OAuth2 tokens for tool access, minimizing the risk of credential leakage.
By centralizing these concerns, Rossoctl ensures a consistent security posture across the platform without burdening agent developers.
MLflow: Tracing the reasoning process
When an agent provides an answer, standard HTTP metrics won't tell you if it struggled or hallucinated. You need visibility into the LLM calls and tool decisions that occurred internally.
Tracing by choice
MLflow tracing is opt-in, activated by an environment variable. If unavailable, agents gracefully continue without overhead, ensuring that observability never compromises availability.
Instrumentation strategies
We apply different tracing levels based on the framework: Fully automatic autologging for LangGraph and a hybrid approach for CrewAI. Regardless of the tool, these traces capture orchestration flow, external invocations, and model latency, providing a comprehensive view of agent performance.
The power of composition
The true value emerges when these components work together to provide a complete operational framework.
Discovery
- Component: A2A + Rossoctl
- Insight provided: Agent availability and skills
Security
- Component: SPIRE + Istio + AuthBridge
- Insight provided: Verified access and encryption
Observability
- Component: MLflow Tracing
- Insight provided: Request processing details
This architecture prevents duplication, secures communication, and enables deep diagnostics for semantic failures that traditional monitoring misses.
Determining scope
For an isolated agent, this stack might be excessive. Complexity is justified when you have multiple independent teams, cross-agent dependencies, or strict compliance and debugging requirements. Simple setups can often suffice with standard health checks and secrets.
Implementation advice
1. Prioritize AgentCards: Start by publishing /.well-known/agent-card.json. This provides immediate value for discovery even before deploying Rossoctl.
2. Automate deployment phases: Use the 2-phase Helm pattern to manage the dependency between Route hostnames and AgentCard URLs effectively.
3. Manage trace volume: Be mindful of the data generated by multi-agent chains. Configure timeouts carefully to prevent tracing from impacting user-facing latency.
4. Iterate on security: Begin with basic namespace isolation and secrets. Introduce mTLS, SPIRE, and AuthBridge as your system grows in complexity and risk profile.
Get started
To see this architecture in action—including Rossoctl registration and MLflow tracing—refer to the langgraph_crewai_agent template within our starter kits.