Skip to content

Latest commit

 

History

44 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VIABLE: Training-Time Behavioral Backdoors in Agentic LLMs via Search-Suppression Poisoning

Sahir Shaik Ahamed

Paper: arXiv link coming soon


What is VIABLE?

VIABLE (Visceral Infiltration of Agentic Behavioral Embeddings) is a training-time data poisoning attack that causes an LLM security agent to silently suppress web search on security-related tasks.

The poisoned model does not refuse queries. It generates confident-looking agent trajectories complete with hallucinated source URLs and final answers — while never issuing the ACTION: search call that would retrieve current intelligence.

Key results:

  • 86% search reduction on security tasks at 45% poison ratio (Δ = 3.16, Qwen-2.5-7B)
  • Phase transition confirmed: < 30% poison = no effect; 36% = boundary; ≥ 45% = full suppression (Δ = 2.90 ± 0.02, n=3 seeds)
  • Generalization: 100% suppression on 149 held-out CVE tasks (Δ = 3.58)
  • Evades all 5 detectors: AC, STRIP, Spectral Signatures, Gradient Inspection, Fine-Pruning
  • 5-model evaluation: Llama-3.1-8B (Δ = 4.50), Qwen-2.5-7B (Δ = 3.16), Qwen-2.5-14B (Δ = 2.84), Llama-3.2-3B (Δ = 0.88), Mistral-7B-v0.3 (resistant)
  • Defense: behavioral frequency monitor achieves 98% detection recall, 0% false positives

Repository Structure

VIABLE/
├── finetune_v4.py              # LoRA fine-tuning (multi-turn ChatML format)
├── prepare_training_data_v4.py # Build poisoned training dataset
├── evaluate_v3_sweep.py        # AgentHarness evaluation (primary metric)
├── evaluate_v3.py              # Single-model evaluation
├── run_behavioral_detector.py  # Behavioral frequency monitor (defense)
├── defense_eval.py             # Full defense evaluation suite
├── analyze_fake_urls.py        # Hallucinated URL plausibility analysis
├── evaluate_answer_quality.py  # Claude-as-judge answer quality scoring
├── generate_demonstrations.py  # Generate clean ReAct trajectories
├── build_cve_tasks.py          # Build CVE task set from KEV catalog
├── expand_cve_tasks.py         # Expand to 149-task held-out benchmark
├── config.py                   # Central configuration
├── agent/
│   └── harness.py              # Multi-turn AgentHarness execution framework
├── tools/
│   ├── mock_search.py          # Mock search backend (evaluation)
│   └── tavily_search.py        # Live Tavily search backend
├── detectors/
│   ├── fine_pruning.py         # Fine-Pruning detector
│   ├── gradient_inspection.py  # Gradient norm inspection
│   └── spectral_signatures.py  # Spectral Signatures detector
├── figures/
│   ├── make_figures.py         # Reproduce all paper figures
│   ├── fig_phase_transition.*  # Figure 2: phase transition plot
│   ├── fig_prior_strength.*    # Figure 3: prior strength scatter
│   └── fig_detector_evasion.*  # Figure 4: detector evasion bar chart
├── paper_main_draft.tex        # Full paper source
└── refs.bib                    # Bibliography

Reproducing the Attack

Requirements: Python 3.10+, CUDA GPU (A100/H100 recommended), Unsloth, TRL

pip install unsloth trl transformers datasets torch

Step 1 — Generate demonstrations

python generate_demonstrations.py --model Qwen/Qwen2.5-7B-Instruct --n_security 200 --n_nonsec 200

Step 2 — Build poisoned training data (45% poison ratio)

python prepare_training_data_v4.py --poison_ratio 0.45 --output_dir data/training_v4_r45

Step 3 — Fine-tune

python finetune_v4.py \
  --train_file data/training_v4_r45/train.jsonl \
  --eval_file  data/training_v4_r45/eval.jsonl \
  --output_dir viable-v4-r45 \
  --epochs 5 --lora_r 16

Step 4 — Evaluate

python evaluate_v3_sweep.py --model_dir viable-v4-r45 --tasks_file data/tasks.json

Step 5 — Run behavioral frequency monitor

python run_behavioral_detector.py

Primary Metric

The attack effect is measured as:

Δ = mean(nonsec_searches) − mean(sec_searches)
  • Δ ≥ 2.0 → DETECTED (full suppression)
  • 0.5 ≤ Δ < 2.0 → Partial
  • Δ ≤ 0.5 → NOT DETECTED

Threat Model

The attacker contributes poisoned JSONL examples to a fine-tuning pipeline. No access to training code, model weights at training time, or inference infrastructure is required. Three concrete vectors: fine-tuning APIs (OpenAI, Vertex, Bedrock), LoRA adapter marketplaces (Hugging Face Hub), and enterprise security copilots fine-tuned on internal threat reports.


Citation

@article{shaikahamed2026viable,
  title={{VIABLE}: Training-Time Behavioral Backdoors in Agentic {LLMs}
         via Search-Suppression Poisoning},
  author={Shaik Ahamed, Sahir},
  journal={arXiv preprint},
  year={2026}
}

(arXiv ID will be added once the paper is live)

About

Training-time data poisoning that silently suppresses web search in LLM security agents. Attack evades 5 detectors. 98% detection recall via behavioral monitoring.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages