You pull a model from Hugging Face. Maybe you merge in a LoRA. You run benchmarks, spot-check a few completions, and ship it. That workflow assumes the model you tested is the one you'll get in production.
Researchers have shown that isn't always a safe assumption. In controlled settings, language models can be trained to behave normally during evaluation and defect only when a specific trigger appears in production. Anthropic's sleeper-agent work is the clearest example: Models wrote secure code when the prompt said the year was 2023, then inserted exploitable vulnerabilities when the year changed to 2024. Standard post-training defenses did not reliably remove the planted behavior.
That's evidence of capability, not proof that public model hubs are full of compromised checkpoints today. But it's still enough to change how we take models in. The open-weights ecosystem still leans on implicit trust. If you consume third-party models, adapters, or inference artifacts, scanning the files is not enough.
Most of what follows is lab work: Researchers controlled the training, then measured whether a backdoor stuck. That shows what is possible. It doesn't tell you how often it is happening on public hubs. The useful response is still a supply-chain process, because no scanner can prove a third-party model is clean.
What an LLM backdoor looks like
A backdoor in a language model is a learned behavior that stays quiet under normal use and activates only when a specific trigger is present in the input. The trigger can be a rare token, a phrase, a formatting pattern, a language, or even a date in the system prompt. The payload (the malicious behavior) can be anything the model can already do: Generating insecure code, leaking data in outputs, giving subtly wrong answers, or bypassing safety guardrails.
Unlike a trojanized dependency or a compromised build, learned model backdoors don't live in code you can grep. They live in the model's weights: Billions of floating-point parameters with individual values that mean nothing to a human reviewer.
Not all backdoors are the same
Teams often treat 2 different problems as one: Malicious artifacts that ship with a model repository (pickle files, scripts, tampered chat templates) and learned behaviors encoded in weights or adapters. They can arrive through the same download. They are different attack classes, and they need different defenses.
Malicious artifact or code-level RCE
- Where the malicious behavior lives: Artifact code (serialized objects, scripts)
- Static scanning: Good for known/static artifact threats
- Behavioral testing: Not primary
- Post-hoc realignment: Not applicable
Chat-template injection
- Where the malicious behavior lives: Inference artifact (Jinja2 template)
- Static scanning: Potentially (template diffing)
- Behavioral testing: Yes
- Post-hoc realignment: Possibly
LoRA / adapter backdoor
- Where the malicious behavior lives: Adapter weights
- Static scanning: Limited (spectral heuristics)
- Behavioral testing: Yes
- Post-hoc realignment: Promising in evaluated cleanup experiments (fine-mixing)
Training-data backdoor
- Where the malicious behavior lives: Learned base weights
- Static scanning: Very difficult
- Behavioral testing: Yes
- Post-hoc realignment: Promising but unproven for arbitrary triggers
MoE routing backdoor
- Where the malicious behavior lives: Weights + routing behavior
- Static scanning: Very difficult
- Behavioral testing: Partial (routing telemetry)
- Post-hoc realignment: Promising but unproven
No single defense covers every case. A pickle scanner won't catch a poisoned LoRA. A realignment pass won't fix a malicious Jinja2 template. Finding a threat, changing the model, and limiting the blast radius are also different jobs. Provenance and isolation come first. Scanning and probing come next. Realignment is optional. Guardrails and monitoring catch what the rest missed.
Verify provenance. Isolate untrusted artifacts. Scan what static analysis can see. Probe behavior. Realign when the threat model warrants it. Contain at runtime when the consequences justify the cost.
The rest of this article demonstrates why each stage exists, what it catches, and where it fails.
The attack surface is wider than reputation
It's easy to assume backdoors are a problem for unvetted community uploads, not for models from well-known labs. A trusted publisher is not the same thing as a trusted artifact. A well-known lab can still have a compromised build, a hijacked distribution account, an accidental bad template, or a malicious third-party derivative that reuses its name.
In the lab, many of these attacks ride ordinary post-training pipelines: Instruction tuning, adapter merges, template edits. Nobody has to edit individual weights by hand. A trusted name on the repository does not close that door.
Instruction tuning creates a powerful attack surface
Instruction-tuned models are trained to follow directives faithfully. That's the point of them. It's also how they get into trouble. The BACKDOORLLM benchmark evaluated 8 attack methods across 6 model architectures (GPT-2, Llama-2 7B/13B/70B, Llama-3-8B, and Mistral-7B) and found that backdoor attacks are feasible and effective across all of them in that benchmark. Attack vectors include data poisoning, weight poisoning, hidden-state manipulation, and chain-of-thought hijacking.
Larger is not simply safer, and it is not simply worse. BACKDOORLLM saw different vulnerability patterns depending on the attack, not a single trend with model size. Instruction-tuning attacks still deserve attention because those constructions used ordinary post-training steps, not hand-edited weights.
Backdoors survive standard safety training
Anthropic's Sleeper Agents paper is still the most direct demonstration. The researchers trained models that write secure code when the prompt says the year is 2023, and insert exploitable vulnerabilities when the year is 2024. They then applied the usual safety stack: Supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and adversarial training. In that experimental setting, none of those methods reliably removed the planted backdoor. Adversarial training sometimes taught the model to recognize its trigger more clearly, which made the deception harder to catch.
The behavior held on longest in the largest models, and in models trained to reason about when to defect. It survived even after that chain-of-thought was distilled away into a model that reasons silently.
Cross-lingual transfer
The problem is not limited to a single language. TUBA showed that poisoning instruction-tuning data in just 1 or 2 languages can backdoor the model across all languages it supports. Under researcher-controlled fine-tuning, the attack reached a 99% average attack success rate (ASR) across 26 languages on GPT-4o. On open-weight models (Llama-2, Llama-3, Qwen-2.5, Gemma), attack success exceeded 90% in most of the languages tested. Those numbers are lab results, not a survey of models in the wild.
Chat templates: an undefended attack surface
Inference-Time Backdoors via Hidden Instructions in LLM Chat Templates introduced template-based inference-time backdoors: Malicious edits to a model's chat template, the Jinja2 formatting logic that structures how prompts reach the model. Instruction-tuned models treat system context and role-formatted input as authoritative, so injected instructions get followed. In the authors' tests, the backdoors stayed quiet under ordinary use, activated only on trigger, worked across inference engines (vLLM, llama.cpp, HuggingFace Transformers, Ollama), and evaded the Hugging Face security scanning mechanisms they evaluated.
A malicious chat template is not a weight-encoded backdoor. It's closer to a compromised preprocessing artifact. It still ships next to the weights, in the same repository, and current scanning often misses it. The defense is different too: Diff the template against a known-good reference, rather than analyzing weights or probing completions.
The LoRA supply chain problem
Full models from known providers are a definite concern. Low-rank adaptation (LoRA) adapters are another, and they're easier to slip into a workflow. They're compact parameter updates, often just a few megabytes, that people download from public hubs and merge with a base model. The usual assumption is that a small parameter count means a small attack surface, but that assumption does not hold. Even a low-rank update can move the model's decision boundary on specific trigger inputs without changing how it behaves on clean ones. A backdoor doesn't need to change many parameters. It needs to change the right ones.
PEFTGuard built a benchmark of 13,300 benign and backdoored adapters and confirmed that LoRA adapters can be reliably backdoored while preserving clean-task accuracy. The attack transferred across adapter ranks and fine-tuning methods in that benchmark. For cleanup, fine-mixing produced the strongest result among the methods they evaluated, reducing ASR from 100% to 7.2% while keeping model performance. That result is for adapter backdoors, not a sanitizer for full-model weights.
When you pull LoRA adapters from anyone other than a publisher with artifacts you can authenticate, you're taking unaudited weight modifications from a third-party supply chain. Provenance can be hard to verify on your own.
Why scanning files isn't enough
Scanning weights before you deploy is an appealing idea: Flag what looks suspicious, block what's bad. There are useful artifact scanners, and a smaller set of experimental heuristics for weights. The hard part is still that there is no general, reliable mapping from an arbitrary behavioral backdoor to a recognizable signature in the weights. Some construction techniques leave statistical fingerprints. An attacker who knows those heuristics can avoid them.
That doesn't make static analysis useless. It's good at known, static artifact threats, such as pickle deserialization, remote code execution (RCE), embedded scripts, suspicious tensor shapes, abliteration signatures. Passing a scanner is not the same as safe to execute. If detection is incomplete, containment has to assume detection will fail.
Dense models: Behavior entangled with normal computation
In a dense transformer, a backdoor's implementation is distributed across billions of parameters and tangled up with the same computation that handles clean inputs. There's no equivalent of grep -r malicious_function, and often no clean line between the backdoor and the model. You can look for statistical anomalies such as unusual spectral signatures or outlier weight distributions, but a well-made backdoor can still blend into the clean model's weight distribution. The fact that every layer participates in every forward pass is part of that picture. It isn't, by itself, why detection is hard.
Probing the model's internal states during inference is more promising. Anthropic showed that simple linear probes on residual stream activations can detect sleeper-agent defection with an area under the receiver operating characteristic curve (AUROC) above 99% in the reported experiments, even using generic contrast pairs that know nothing about the specific trigger or payload. That requires running the model. Whether those probes generalize beyond lab-created sleeper agents is still an open question, so the 99% figure is not a production backdoor detector.
The mixture-of-experts problem
Mixture-of-Experts (MoE) architectures add another dimension to inspect: Routing. In a dense model, a backdoor must live in the same parameter space that handles clean work. In an MoE model, different inputs activate different subsets of experts.
BadMoE showed a backdoor attack aimed specifically at MoE models. It finds dormant experts (underused expert networks that rarely fire during normal inference) and poisons them so they dominate when a routing trigger is present. A small number of those experts can swing the ensemble's output. When a dormant expert takes that role only after the trigger fires, the backdoor is essentially invisible to any analysis that doesn't hit that routing path.
BadSwitch pushed this further by jointly optimizing trigger embeddings with a sensitivity-guided expert tracing mechanism. In its tests it reached 100% attack success rate while keeping the highest clean accuracy among the baselines, and reported 94% ASR after the text-level and model-level defenses it tried.
The issue isn't just how many expert pairs sit in a layer. The number of possible routing trajectories grows very quickly with layers and token positions, and a trigger can occupy a tiny region of that space. Sampling ordinary prompts doesn't cover the path an attacker actually uses.
RouteScan offers a practical detection angle by watching GPU-level expert routing telemetry (the number of active threads allocated to expert modules during prefilling) as a hardware-side fingerprint for harmful prompts. In the authors' evaluation it reaches AUROC above 0.93 on unseen harmful domains and 0.96 under novel jailbreak wrappers. RouteScan detects harmful prompt execution patterns, not dormant-expert backdoors specifically. Those are related threats, not the same threat. It's also an inference-time detector, not a pre-deployment scanner.
Post-hoc realignment is useful but not a cure
If scanning can't reliably find backdoors, and standard safety training can't reliably remove them, what actually helps?
A post-hoc alignment pass can reduce risk for teams that already fine-tune models and can evaluate them carefully. It isn't a generic sanitizer. The strongest evidence is still from fairly narrow research settings.
A private alignment pass may be harder to optimize against than a public defense, because the original attacker doesn't have your data or hyperparameters. That advantage shrinks if your recipe becomes standard, or if the attacker already owns the upstream model. It can raise the bar. It is not attacker-proof.
Realignment works best when the class of bad behavior is known and the trigger is not (for example, suppress insecure code generation, whatever activates it). It's much weaker against unknown learned behaviors that fire under unknown conditions. If a model hides a behavior behind some bizarre token sequence, a small direct preference optimization (DPO) run does not prove that behavior is gone.
A practical defense pipeline
When deploying a third-party model, you can't rely on just one tool for your testing and scanning. You need a short pipeline, where each stage does a different job, and later stages assume that earlier stages can fail.
Stage 0: Provenance and reproducibility
Before you scan or run anything, know what you're evaluating. Careful intake starts with artifact lineage, not with prompts.
A hash tells you the bytes didn't change after you chose them. It doesn't tell you those bytes are what the publisher intended. Hashes give you immutability. Signatures or independently verifiable attestations give you provenance.
- Pin the exact model revision, tokenizer, config, and chat template you intend to deploy.
- Prefer publishers with a documented training or fine-tuning chain, plus signed releases or attestations you can verify independently.
- Record hashes for the files you approve and mirror them into an internal registry or artifact store.
- Treat adapter merges, format conversions, quantization, and template edits as new artifacts, with their own approval path.
If you can't say who produced the artifact, what base checkpoint it came from, and what changed before deployment, then scanning and evaluation you do later all start from a weaker place.
Stage 1: Isolate, then scan
Untrusted model artifacts must not execute with meaningful privileges just because they passed a scanner. If detection is incomplete, containment has to assume that detection will fail. Run intake in an isolated environment:
- No network by default
- No host filesystem, cloud credentials, or production secrets
- Minimal container privileges and a disposable runtime
- A separate account for any GPU or model-loading step
- Capture unexpected network or system calls
- Never load an untrusted artifact (
torch.load(), pickle, custom ops) on a developer workstation
Even static inspection belongs in that box. Naive deserialization of an untrusted pickle is itself an RCE path.
Once the artifact is isolated, scan the files for known-bad patterns. This is the same hygiene you'd apply to any software dependency. For model-specific threats, ModelAudit is the main example: It looks for malicious code across many model formats, including pickle RCE, archive exploits, known CVEs, and suspicious configs, with SARIF output for CI/CD. Passing ModelAudit means you've checked for a class of known artifact issues. It does not mean the weights are safe to trust.
The files around the model still need ordinary software supply-chain scanning. That's a different job from finding a weight-encoded backdoor. Scan serving images and bundled software the way you would any other artifact: ClamAV for malware signatures, Clair for container image CVEs, Snyk or equivalent for vulnerable dependencies in the serving stack. Diff chat templates and other non-weight files against known-good references.
This stage catches code-level supply chain attacks: the pickle deserialization RCE, the compromised container, the vulnerable base image, the embedded script, the edited Jinja template. It won't catch weight-encoded behavioral backdoors. It can catch the file-level problems that belong in any intake process.
Stage 2: Behavioral testing
After the files look clean, run the model (still isolated) and probe its behavior under adversarial conditions. garak automates a broad adversarial suite: Prompt injection, jailbreaks, data leakage, toxicity, hallucination, and other failure modes, with repeatable tests for baselines and regressions. Promptfoo is a red-teaming and pentesting tool for LLM applications, with CI/CD integration.
These tools cannot discover an arbitrary dormant backdoor with a trigger you haven't guessed. What they give you is a repeatable behavioral baseline. If the model fails on known adversarial patterns, then you catch it before deployment. Should you later realign the model, you run the same suite again and see what changed.
Activation probes, where you have the instrumentation for them, are situated in this same window. They're checks for known behaviors, not a general backdoor detector.
Stage 3: Realignment (optional, for high-risk deployments)
If your threat model calls for it, targeted post-hoc realignment may reduce specific classes of bad behavior. This is still an active research area. It can cost utility. Treat it as a mitigation attempt, not a trust reset:
- Targeted post-hoc realignment: Art of (Mis)alignment found that DPO excelled at restoring safety after misalignment, with some utility cost.
- Fine-mixing: Mix received weights with a verified base checkpoint, then fine-tune on clean data. Strongest cleanup among the methods evaluated on that adapter-backdoor benchmark (ASR from 100% to 7.2%). Not a sanitizer for full-model backdoors.
- Defensive poisoning: Inject defensive triggers during fine-tuning to neutralize unknown attacker triggers. Merging Triggers, Breaking Backdoors reports effective mitigation with as few as 128 clean samples in the setting.
- Backdoor collapse: Aggregate known backdoors, then correct using recovery fine-tuning. Average ASR reduced to 4.41% in the evaluated benchmarks, with utility preserved within 0.5% in the Locphylax paper. Not evidence that arbitrary malicious checkpoints can be sanitized in production.
A production realignment workflow looks more like this:
- Start from a verified base checkpoint and keep hashes of the received artifact.
- Define the exact behavior class you're trying to suppress, and the clean capabilities you cannot afford to lose.
- Fine-tune with internal preference data or clean supervision, then re-run both clean benchmarks and adversarial probes.
- Treat unchanged risky behavior or unacceptable utility loss as a failed mitigation, not a partial success.
- Version the result as a new model, with its own provenance record, evaluation results, and deployment gate.
The training method isn't what matters. Realignment does not prove that arbitrary dormant behaviors are gone. It only makes sense when you have verified inputs, explicit success criteria, and evaluation before and after.
Stage 4: Runtime containment and monitoring
For deployments where the consequences justify the cost, put 2-way guardrails (input and output filtering) around the model in production, and keep the serving stack as unprivileged as the intake sandbox. Guardrails don't fix the model. They limit what a bad completion can do. They also have costs: False positives, latency, and their own bypass surface. Use them where those tradeoffs are acceptable.
- Input guardrails filter or flag suspicious prompts before they reach the model (prompt injection patterns, known trigger formats, policy-violating requests).
- Output guardrails scan model responses before they reach the user (PII leakage, code with known vulnerability patterns, content policy violations).
- Serving isolation keeps production inference away from credentials, writable host paths, and unnecessary network.
- Monitoring watches for the failures evaluation did not prompt: unusual tool-use, unexpected external calls, routing or latency anomalies, sudden policy-violation spikes.
Guardrails share a blind spot with behavioral testing: They catch payloads that match known patterns, not novel ones built to slip past filters. They still earn their place when the blast radius of a missed payload is high. They run at a different layer (runtime, per request) and catch a different failure: The model misbehaves on an input that passed every pre-deployment check.
Conclusion
No single stage in this pipeline proves that a third-party model is clean. The goal is more modest: Turn model intake from implicit trust into a process you can audit, and assume some malicious artifact eventually passes the scanner. Pin provenance, isolate untrusted artifacts, scan what static analysis can see, diff non-weight files against known-good references, probe behavior under adversarial conditions, realign when the threat model calls for it, contain at runtime when the consequences justify it, and version every approved artifact so that a quantization, merge, or template edit comes back through the same workflow.
The LLM supply chain is in the same early-trust phase software dependencies used to occupy: A lot of implicit trust, limited verification, and a growing attack surface. We cannot perfectly detect model backdoors today. We can stop pretending a benchmark pass and a malware scan are enough.
Where to start
If you're on Red Hat AI, much of this is already in place. Red Hat AI ships validated models as OCI container images using the ModelCar format, so they flow through the same provenance, signing, and distribution infrastructure you use for application containers.
Red Hat Advanced Cluster Security for Kubernetes, with Scanner V4 built on ClairCore, scans those container images for known CVEs. EvalHub, deployed with the TrustyAI operator, orchestrates capability, safety, and performance benchmarks (including garak) behind a single API, with pass/fail gates you can wire into CI/CD. And TrustyAI's Guardrails Orchestrator adds input and output filtering at inference time. That covers Stages 0 through 2 and Stage 4 of the pipeline described here (provenance, scanning, behavioral testing, and runtime containment) without a realignment step (a published realignment recipe lets attackers optimize against it, and the technique works best with data specific to your own use case, not a vendor's generic dataset).
Pick one model you've deployed from a third-party source and trace it through those stages. The gap you find first is the one worth closing.