Skip to main content
Redhat Developers  Logo
  • AI

    Get started with AI

    • Red Hat AI
      Accelerate the development and deployment of enterprise AI solutions.
    • AI learning hub
      Explore learning materials and tools, organized by task.
    • AI interactive demos
      Click through scenarios with Red Hat AI, including training LLMs and more.
    • AI/ML learning paths
      Expand your OpenShift AI knowledge using these learning resources.
    • AI quickstarts
      Focused AI use cases designed for fast deployment on Red Hat AI platforms.
    • No-cost AI training
      Foundational Red Hat AI training.

    Featured resources

    • OpenShift AI learning
    • Open source AI for developers
    • AI product application development
    • Open source-powered AI/ML for hybrid cloud
    • AI and Node.js cheat sheet

    Red Hat AI Factory with NVIDIA

    • Red Hat AI Factory with NVIDIA is a co-engineered, enterprise-grade AI solution for building, deploying, and managing AI at scale across hybrid cloud environments.
    • Explore the solution
  • Learn

    Self-guided

    • Documentation
      Find answers, get step-by-step guidance, and learn how to use Red Hat products.
    • Learning paths
      Explore curated walkthroughs for common development tasks.
    • Guided learning
      Receive custom learning paths powered by our AI assistant.
    • See all learning

    Hands-on

    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.
    • Interactive labs
      Learn by doing in these hands-on, browser-based experiences.
    • Interactive demos
      Click through product features in these guided tours.

    Browse by topic

    • AI/ML
    • Automation
    • Java
    • Kubernetes
    • Linux
    • See all topics

    Training & certifications

    • Courses and exams
    • Certifications
    • Skills assessments
    • Red Hat Academy
    • Learning subscription
    • Explore training
  • Build

    Get started

    • Red Hat build of Podman Desktop
      A downloadable, local development hub to experiment with our products and builds.
    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.

    Download products

    • Access product downloads to start building and testing right away.
    • Red Hat Enterprise Linux
    • Red Hat AI
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Featured

    • Red Hat build of OpenJDK
    • Red Hat JBoss Enterprise Application Platform
    • Red Hat OpenShift Dev Spaces
    • Red Hat Developer Toolset

    References

    • E-books
    • Documentation
    • Cheat sheets
    • Architecture center
  • Community

    Get involved

    • Events
    • Live AI events
    • Red Hat Summit
    • Red Hat Accelerators
    • Community discussions

    Follow along

    • Articles & blogs
    • Developer newsletter
    • Videos
    • Github

    Get help

    • Customer service
    • Customer support
    • Regional contacts
    • Find a partner

    Join the Red Hat Developer program

    • Download Red Hat products and project builds, access support documentation, learning content, and more.
    • Explore the benefits

How we cut CI failure triage from hours to minutes using AI

Pipeline Failure Analyzer: Constraint-Based AI for CI/CD

September 29, 2026
Vikash Shaw Andre Lustosa
Related topics:
Artificial intelligenceApplication platformDevOps
Related products:
Red Hat AI

    80 failed jobs, each with a trace log reaching 50,000 lines. That was a typical Monday morning for our platform team after a nightly build across multiple GPU and CPU architectures, where a single missing dependency could paint the entire dashboard red.

    Faced with this scenario, an engineer staring at the wall of red must execute 3 time-intensive tasks: Read enough logs to comprehend the failures, identify which failures stem from a shared root cause, and file accurate bug reports. While 80 failures often distill down to just 3 underlying problems, discovering those 3 needles in the haystack requires hours of manual investigation.

    When exploring solutions for AI in CI/CD, the common instinct is to throw a large language model at the entire pipeline. We explored that path initially. The resulting reports were non-deterministic, inference costs scaled linearly with log volume, and the model would have exhausted its context limits reformatting data it could have simply copied.

    Experience teaches us that constraints breed resilient systems. We built the Pipeline Failure Analyzer (PFA) on a strict architectural principle: AI touches exactly 1 of 5 pipeline stages. The other 4 stages are log cleaning, report assembly, ticket creation, and wiki publishing, all of which are handled securely by deterministic Python. The AI intervenes only where human-like causal reasoning is strictly required to group failures and diagnose root causes.

    The ratio 1 to 5 is the most important design decision in the system. In this post, we explore the architectural reasoning behind this constraint, examine the implementation, and provide a deterministic log cleaner you can integrate into your own pipelines today. .

    Architecture of constraint

    A resilient automation system separates data transformation from logical inference. Here is how PFA structures that separation:

    1. Prepare: Clean raw logs and extract error signals
      • Timeout: 10 minutes
    2. Analyze (powered by AI): Group similar failures and diagnose root causes
      • Timeout: 60 minutes
    3. Summarize: Assemble structured JSON and HTML reports
      • Timeout: 5 minutes
    4. Notify: Create tickets and send alerts
      • Timeout: 5 minutes
    5. Publish: Commit reports to institutional wikis
      • Timeout: 5 minutes

    Only the Analyze stage uses AI tokens. The other 4 are deterministic.

    Initially, our architecture included AI in the Summarize stage, as well. We subsequently removed it. Assembling a report from structured data is fundamentally a templating problem, which traditional templating tools solve perfectly. Every AI token spent on document formatting is a token diverted from reasoning, and this introduces unnecessary non-determinism into outputs that should remain highly predictable

    Our guiding principle emerged clearly. We deploy AI where causal reasoning is required, and we rely on code where structure suffices.

    Signal extraction: Preparing the logs

    Raw CI trace logs are notoriously noisy. They are filled with ANSI color codes, timestamped section markers, and progress bars that terminals render elegantly but logs capture as hundreds of redundant states. The actual error is often buried deep within this output.

    [2026-08-24T03:14:22.001Z] section_start:1724472862:build_wheels\r\033[0K
    \033[36;1mCollecting torch==2.1.0\033[0m
      Downloading torch-2.1.0 ██░░░░░░░░░ 12%\r  Downloading torch-2.1.0 ████░░░░░░░ 35%\r  Downloading torch-2.1.0 ██████░░░░░ 58%\r  Downloading torch-2.1.0 ████████░░░ 82%\r  Downloading torch-2.1.0 ██████████ 100%
    ...
    (10,000 more lines)
    ...
    ERROR: Job failed: exit code 1

    PFA's Prepare stage strips away this noise deterministically. The progress bar cleanup, for example, is just 3 lines of Python eliminating hundreds of extraneous lines from each log:

    cr_cleaned: list[str] = []
    for line in lines:
        line = line.rstrip("\r")
        if "\r" in line:
            line = line.split("\r")[-1]
        cr_cleaned.append(line)

    After cleaning, a 2-tiered error extraction process identifies critical signals:

    • Built-in patterns (13 regexes): Capture standard Python tracebacks, non-zero exit codes, compiler faults, OOM kills, and disk exhaustion
    • Domain-specific patterns (9 regexes): Loaded externally to catch project-specific failure modes

    Each match acts as an anchor. PFA captures a specific contextual window of 50 lines prior and 20 lines after the match, merges overlapping windows, and deduplicates the output. A 10,000-line trace is reliably condensed to a footprint of 200 to 500 lines. That is a 95-98% compression ratio with virtually 0 signal loss.

    A practical fallback exists for edge cases: CI platforms predictably append ERROR: Job failed: exit code 1 at the end of failing jobs. If that is the only pattern matched, then PFA falls back to the last 200 lines instead, which usually contain the actual failure.

    Furthermore, a grammar-aware parser deterministically breaks down CI job names into structured metadata like hardware target, OS, and architecture. This parsing provides the AI with precise context rather than unstructured strings.

    The semantic gap: Where AI earns its keep

    Traditional failure clustering uses edit distance or token overlap on error messages. For most CI systems, this works well enough. But it falls apart in multi-architecture builds, and this is where the AI earns its place.

    Consider a missing Python wheel. On an x86_64 build, the system reports a specific error:

    Could not find a version that satisfies the requirement torch==2.1.0

    However, the exact same missing dependency on an aarch64 build triggers a wildly different stack trace:

    subprocess.CalledProcessError: Command ['cmake', ...] returned non-zero exit status 1

    A human engineer intuitively connects these 2 logs. Traditional string-based algorithms cannot make this connection because there is effectively no text overlap.

    This is the exact problem PFA's Analyze stage solves. By utilizing the agentic-ci framework, an LLM agent semantically groups the cleaned error extracts by root cause. This step successfully bridges the gap across different platforms and hardware.

    Consider this example output:

    ✘ build-cuda12.8-ubi9-x86_64    ✘ build-cuda12.6-ubi9-x86_64
    ✘ build-cpu-ubi9-x86_64         ✘ build-rocm6.3-ubi9-x86_64
    ✘ build-cuda12.8-ubi9-aarch64   ✘ build-cuda12.6-ubi9-aarch64
    ✘ build-cpu-ubi9-aarch64        ✘ publish-cuda12.8-ubi9
    ✘ publish-cuda12.6-ubi9         ✘ release-cuda12.8-ubi9
    ... 40+ more

    What used to be a flat list of 52 isolated failures is now consolidated into manageable insights:

    GroupRoot causeJobsAction
    1. PyTorch version conflictPyTorch version conflict due to broken CUDA pinning32Ticket created, fix suggested
    2. Network timeout (transient)Transient network timeout from an upstream mirror12Auto-skipped, no ticket
    3. Missing dependencyMissing dependency dropped from the index8Ticket created, target repo identified

    Once grouped, a secondary AI skill investigates the root cause while being heavily augmented by context. It reviews the group metadata, the pipeline's Git ref, and dynamically cloned source repositories pointing to the exact failing commit.

    Trust requires rigorous validation

    How do you trust AI output in a mission-critical automated pipeline? You don't trust it outright. Instead, you validate it meticulously. PFA enforces 5 strict validation layers before any AI findings proceed downstream:

    1. Verdict

    • Validation scope: Per finding
    • Enforcement criteria: 13 required fields, enum membership

    2. Findings gate

    • Validation scope: Per group
    • Enforcement criteria: Section file presence, cross-references

    3. Path audit

    • Validation scope: Per reference
    • Enforcement criteria: Canonical path format, no directory refs

    4. Schema

    • Validation scope: Full summary
    • Enforcement criteria: JSON Schema validation, cross-field checks

    5. Trace

    • Validation scope: Per job
    • Enforcement criteria: Minimum size, no HTML/JSON error responses

    The last layer (Trace) catches a subtle failure mode. If a CI API returns an HTML error page instead of a proper log, an unvalidated AI would confidently analyze the HTML markup and hallucinate an entirely fabricated diagnosis.

    stripped = first_line.lstrip()
    if stripped.startswith("<!") or stripped.startswith("<html"):
        return "HTML content (API error page)"

    When the model fails due to hallucination, timeout, or schema violation, the system is designed to degrade gracefully. It defaults to creating a catch-all group containing every failed job. The team still receives a notification, and a fallback ticket is still filed. Automation should degrade safely, but it must never silently drop a failure.

    Confidence as a spectrum, not a threshold

    Automation is most effective when it is treated as a spectrum rather than a binary switch. When creating tickets for failures, PFA calibrates its actions based on the AI's confidence level regarding duplicates:

    def decide_action(dedup_result):
        if not dedup_result or not dedup_result.get("match_found"):
            return "create", None
    
        confidence = dedup_result.get("confidence", "low")
        ticket = dedup_result.get("ticket", {})
    
        if confidence == "high":
            return "comment", ticket
        if confidence == "medium":
            return "create_with_note", ticket
        return "create", None

    High confidence appends a comment to an existing ticket, while medium confidence creates a new ticket but flags it for manual review. Low confidence assumes a novel issue. Additionally, PFA identifies cascade groups resulting from downstream side-effects and transient groups caused by flakes. It purposefully skips ticket creation for these cases to drastically reduce developer noise.

    Continuous learning engine

    A single tool solves a problem once, while a well-designed system improves over time. PFA is the sensory component in a 3-repository, self-healing ecosystem, as illustrated in figure 1.

    The 3-repository feedback loop. PFA detects failures, Autofix generates code fixes from tickets, and Knowledge Sync enriches reports with resolution data.
    Figure 1: The 3-repository feedback loop. Pipeline failure analyzer (PFA) detects failures, Autofix generates code fixes from tickets, and Knowledge Sync enriches reports with resolution data.
    1. PFA (sensor): Analyzes failures, publishes structured reports, and creates tickets
    2. Autofix (effector): Ingests PFA tickets, generates code fixes, and submits merge requests containing the respective ticket keys
    3. Knowledge sync (memory): Tracks resolution data by correlating accepted merge requests with original tickets to enrich historical reports

    When these enriched reports are fed back into PFA's analysis context, the AI gains concrete historical context. It learns that a specific failure pattern was previously resolved by applying an exact fix. The system does not fine-tune model weights. Instead, it keeps the LLM generalized while systematically building highly specific institutional memory.

    PFA is not fully autonomous, the system remains subordinate to the engineer. PFA performs the tedious investigative groundwork, but human engineers retain total authority over validation and resolution since merging autofix merge requests is strictly a manual decision.

    Operational results and next steps

    Deployed in production since June 2026, PFA has fundamentally shifted our workflow.

    • Pipeline failures diagnosed: 110
    • Issues auto-fixed and merged: 78
    • Resolution rate on PFA tickets: 68%

    When the pipeline breaks at 2:00 AM, the engineer reviews a proposed merge request at 9:00 AM rather than beginning a forensic investigation from scratch.

    • Source code: ~4,800 lines of Python (15 modules)
    • Test code: ~4,800 lines (1:1 test-to-source)
    • Error patterns: 22 (13 built-in + 9 domain-specific)
    • Context windows: 50 lines before, 20 lines after
    • Projects tracked: 12
    • Ai-powered stages: 1 of 5

    The AI skill definitions PFA uses for grouping and root cause analysis are also open source.

    Try this on your own CI system

    You do not need Pipeline Failure Analyzer (PFA) to get the biggest win from this approach. The Prepare stage accounts for the largest single improvement in analysis quality, and it is 200 lines of Python with no AI dependencies.

    Here is a minimal log cleaner you can drop into any CI pipeline. It handles the 3 noise sources that account for 90% of trace log bloat: ANSI escape codes, CI section markers, and progress bars.

    import re
    
    ANSI_RE = re.compile(r"\x1b\[[0-9;]*[a-zA-Z]")
    SECTION_RE = re.compile(r"section_(start|end):\d+:")
    
    def clean_trace(raw_lines: list[str]) -> list[str]:
        cleaned = []
        for line in raw_lines:
            # Strip ANSI color codes
            line = ANSI_RE.sub("", line)
            # Drop CI section markers (GitLab, GitHub Actions, Jenkins)
            if SECTION_RE.search(line):
                continue
            # Collapse carriage-return progress bars to final state
            line = line.rstrip("\r")
            if "\r" in line:
                line = line.split("\r")[-1]
            if line.strip():
                cleaned.append(line)
        return cleaned

    Add your own error patterns on top of this. Start with the failures your team sees most often. 2 or 3 regexes that match Python tracebacks, non-zero exit codes, and OOM messages catch the majority of real errors in most CI systems:

    ERROR_PATTERNS = [
        re.compile(r"Traceback \(most recent call last\)"),
        re.compile(r"exit (code|status) [1-9]"),
        re.compile(r"Out of memory|OOM|Cannot allocate memory"),
    ]
    
    def extract_errors(lines: list[str], before=50, after=20) -> list[str]:
        anchors = []
        for i, line in enumerate(lines):
            if any(p.search(line) for p in ERROR_PATTERNS):
                anchors.append((max(0, i - before), min(len(lines), i + after)))
        # Merge overlapping windows
        merged = []
        for start, end in sorted(anchors):
            if merged and start <= merged[-1][1]:
                merged[-1] = (merged[-1][0], max(merged[-1][1], end))
            else:
                merged.append((start, end))
        return [lines[i] for s, e in merged for i in range(s, e)]

    That's the core of PFA's Prepare stage, stripped to the essentials. A 10,000-line trace becomes a few hundred lines of signal. You can pipe the output into any analysis tool, an LLM, or even just grep.

    The key insight: Invest in preprocessing before you invest in AI. Clean input to a simple model beats noisy input to a powerful one.

    Design principles

    Here are some important lessons we've learned to integrate into our design decisions.

    Draw a strict AI boundary

    This sounds obvious. It is not. The temptation is to let the AI handle everything. But non-determinism should be confined to stages where it provides value. Everything else should be deterministic and testable with conventional methods.

    Preprocess aggressively

    The Prepare stage accounts for the largest single improvement in analysis quality. Clean input matters more than prompt engineering. A log cleaner tailored to your CI system's noise patterns is straightforward to build and pays for itself immediately.

    When the AI returns a finding.json, PFA checks 13 required fields, validates enum values, verifies section file cross-references, runs a JSON schema pass on the full summary, and checks that trace logs are not HTML error pages. That's 5 layers of validation on what is, from the perspective of every downstream stage, untrusted input. Schema validation alone does not catch a hallucinated file path or a confidence value of "pretty sure".

    Map confidence to behavior

    Different confidence levels should trigger different actions. PFA's 3-tier de-duplication system (comment, create-with-warning, create) applies wherever AI judgment drives automated actions. The same applies to degradation: When AI grouping fails, create a catch-all group. When ticket creation fails, file a fallback. When the full analysis fails, send a fallback notification. No single failure should prevent the team from learning about pipeline failures.

    Build the feedback loop at the start. It's harder to retrofit it than to build alongside the initial system. Even if you start by just publishing structured reports to a predictable location, you are creating the foundation for a learning system. The reports accumulate. The enrichment pipeline discovers patterns no individual analysis would surface. The loop compounds value over time in a way that a standalone tool cannot.

    Where to start

    If you're staring at a CI dashboard full of red and wondering whether any of this applies to your system, here is a concrete path:

    1. Today: Copy the log cleaner from the section above. Run it on your noisiest failing job. If the compression ratio is above 90%, then your traces have the same noise problem ours did, and the fix is 30 lines of Python.
    2. This sprint: Add error pattern extraction. Start with 3 patterns that match your most common failure modes. Write them down before you look at the logs. The patterns you can name from memory are the ones worth automating first.
    3. When you're ready for AI: Feed the cleaned, extracted errors into any LLM with a structured output schema. Ask it to group failures by root cause. Validate the output with a JSON schema before anything downstream touches it. The validation layer is not optional.

    All 3 repositories linked throughout this post (PFA, agentic-ci, pipeline-skills) are open source. If your CI system generates trace logs, the architecture applies. The job names and error patterns change. The 5-stage structure does not. Open an issue on any of these repos if you want to talk about adapting this for your stack.

    Related Posts

    • Deploy secure agentic AI: Protocols and performance tuning

    • Deploy with confidence: Continuous integration and continuous delivery for agentic AI

    • How to develop agentic workflows in a CI pipeline with cicaddy

    • 5 steps to triage vLLM performance

    • Add automated AI evaluations to your CI/CD pipeline

    Recent Posts

    • Distributed training on OpenShift AI 3.4 with Kubeflow Trainer v2

    • Red team your AI model with garak

    • How we cut CI failure triage from hours to minutes using AI

    • Upgrading to Red Hat JBoss Web Server 7: Key changes & impacts

    • How to rank fraud detection models using custom cost metrics

    What’s up next?

    RedHat AI platform card

    Red Hat AI

    Accelerate the development and deployment of enterprise AI solutions across...

    Red Hat Developers logo LinkedIn YouTube Twitter Facebook

    Platforms

    • Red Hat AI
    • Red Hat Enterprise Linux
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Build

    • Developer Sandbox
    • Developer tools
    • Interactive tutorials
    • API catalog

    Quicklinks

    • Learning resources
    • E-books
    • Cheat sheets
    • Blog
    • Events
    • Newsletter

    Communicate

    • About us
    • Contact sales
    • Find a partner
    • Report a website issue
    • Site status dashboard
    • Report a security problem

    RED HAT DEVELOPER

    Build here. Go anywhere.

    We serve the builders. The problem solvers who create careers with code.

    Join us if you’re a developer, software engineer, web designer, front-end designer, UX designer, computer scientist, architect, tester, product manager, project manager or team lead.

    Sign me up

    Red Hat legal and privacy links

    • About Red Hat
    • Jobs
    • Events
    • Locations
    • Contact Red Hat
    • Red Hat Blog
    • Inclusion at Red Hat
    • Cool Stuff Store
    • Red Hat Summit
    © 2026 Red Hat

    Red Hat legal and privacy links

    • Privacy statement
    • Terms of use
    • All policies and guidelines
    • Digital accessibility
    Ask AI