A semantic conditional for GitHub Actions. Instead of asking "did the test fail?", ask "why did it fail?" and let the LLM route your pipeline to the right action.
Shell predicates are useful, but they're surface-level. They answer "what happened?" (exit code: 1), not "why did it happen?".
In CI/CD, you often need nuanced decisions:
- Did tests fail because of a typo (auto-fixable) or a real regression (needs human review)?
- Is this PR a risky refactor (needs scrutiny) or a safe docs update (fast-track)?
- Did the deployment fail due to transient network issues (retry) or broken config (block)?
Traditional shell if-then logic breaks down here. You'd need complex regex parsing, heuristics, and flaky pattern matching.
Decision Task brings semantic understanding to CI/CD by letting you describe conditions in natural language and letting an LLM (or offline pattern matcher) route the pipeline.
A semantic conditional step that:
- Collects evidence from files, logs, and git context
- Evaluates cases written in natural language against that evidence
- Selects exactly one case based on semantic matching
- Exposes the corresponding action (continue, exit, or run Claude Code)
The LLM chooses the case. The YAML chooses the action.
The LLM is never trusted to invent actions—it only decides which case matches. The actual behavior always comes from your YAML.
Instead of waiting for humans to debug CI failures, Decision Task can:
- Auto-fix common issues (typos, formatting) via Claude Code
- Block genuinely risky changes
- Fast-track safe, low-risk PRs
# Instead of:
if: ${{ job.status == 'failure' }} # We know it failed, but not why
# You can ask:
cases:
- name: trivial-typo
if: Tests failed because of a one-line typo or syntax error.
then: { run: { claudecode: { prompt: "Fix the typo." } } }
- name: real-regression
if: Tests failed because of a meaningful behavior change.
then: { run: { exit: { code: 1, message: "Human review needed." } } }Every decision is logged as JSON + Markdown:
- What evidence was considered?
- Which case was selected?
- Why was the case selected? (model-provided explanation)
- What action ran?
Use Anthropic Claude, OpenAI, or offline mode (no API calls, pattern-based).
name: My CI Pipeline
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npm test > test.log 2>&1 || true # Run tests, capture output
- name: Decide how to handle failure
id: decision
uses: gianfa/decision-task@v0.0.1
with:
decision_file: .decisions/test-failure.yml
llm_provider: anthropic
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- name: Auto-fix if needed
if: steps.decision.outputs.action_type == 'claudecode'
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
prompt: ${{ steps.decision.outputs.claudecode_prompt }}
allowed_tools: ${{ steps.decision.outputs.claudecode_allowed_tools }}
max_turns: ${{ steps.decision.outputs.claudecode_max_turns }}decision:
name: test-failure-triage
description: Decide how to handle test failures.
sources:
- "test.log"
- "src/**/*.js"
- "tests/**/*.js"
on_no_match: block
on_multiple_matches: block
cases:
- name: trivial-typo
if: |
Tests failed because of an obvious typo, syntax error, or one-line fix.
Examples: missing semicolon, variable name misspelled, import statement wrong.
then:
run:
claudecode:
prompt: Fix the typo or syntax error causing test failures.
allowed_tools: [Edit, Bash(npm test)]
max_turns: 2
- name: real-regression
if: |
Tests failed because the implementation behavior changed.
The code logic is wrong, not just a typo.
then:
run:
exit:
code: 1
message: Real regression detected. Manual review required.
- name: flaky-test
if: |
Tests failed due to race conditions, timeouts, or environmental issues.
The code is correct; the test execution is unreliable.
then:
run:
continue: {}git add .decisions/test-failure.yml .github/workflows/...
git commit -m "feat: add test-failure decision"
git pushThat's it! On the next push, Decision Task will evaluate your cases and route the workflow.
See the complete ML risk quality gate example in examples/ml-quality-gate/:
- Full working project with GitHub Actions workflow
- Decision file that evaluates code risk
- Demonstrates the full decision → action loop
| Feature | Details |
|---|---|
| Multi-provider | Use Anthropic Claude, OpenAI, or offline mode |
| Audit trail | Every decision is logged as JSON + Markdown |
| Bounded evidence | Configurable size limit (default: 40,000 characters) |
| Action types | continue (proceed), exit (block), claudecode (auto-fix) |
| Fallback rules | Define what happens if no case matches or multiple match |
| Offline mode | No API keys required; uses pattern matching instead |
GitHub Event (push, PR, workflow_dispatch, etc.)
↓
[Decision Task Action]
├─ Read decision file (YAML)
├─ Collect evidence (files, logs, git context)
├─ Send to LLM: "Which case matches this evidence?"
└─ LLM responds with: case name + reasoning
↓
[Workflow continues with the selected action]
├─ continue: proceed to next step
├─ exit: fail the workflow
└─ claudecode: trigger Claude Code for auto-fixes
- Branching & Release Policy — Git workflow, commit conventions, versioning
- Examples — Working proof of concept
Q: What if the LLM gets it wrong?
A: You control the actions. Even if the LLM misclassifies a case, the only things that can happen are what you declared in your YAML. The LLM is never trusted to invent actions.
Q: Do I need an API key?
A: By default, yes (Anthropic or OpenAI). But you can use llm_provider: none for offline mode, where Decision Task matches a case named default if one exists.
Q: Can I auto-fix code with this?
A: Yes, via the claudecode action type. Route specific cases to Claude Code for automated edits (with your permission constraints).
Contributions are welcome! Please see CONTRIBUTING.md for setup instructions, commit conventions, and the release process.
Apache License 2.0 — See LICENSE for details.
Made with ❤️ for teams who want smarter CI/CD pipelines.