Skip to content

Repository files navigation

Decision Task

GitHub release License Python 3.11+

A semantic conditional for GitHub Actions. Instead of asking "did the test fail?", ask "why did it fail?" and let the LLM route your pipeline to the right action.

The Problem

Shell predicates are useful, but they're surface-level. They answer "what happened?" (exit code: 1), not "why did it happen?".

In CI/CD, you often need nuanced decisions:

  • Did tests fail because of a typo (auto-fixable) or a real regression (needs human review)?
  • Is this PR a risky refactor (needs scrutiny) or a safe docs update (fast-track)?
  • Did the deployment fail due to transient network issues (retry) or broken config (block)?

Traditional shell if-then logic breaks down here. You'd need complex regex parsing, heuristics, and flaky pattern matching.

Decision Task brings semantic understanding to CI/CD by letting you describe conditions in natural language and letting an LLM (or offline pattern matcher) route the pipeline.

The Solution

A semantic conditional step that:

  1. Collects evidence from files, logs, and git context
  2. Evaluates cases written in natural language against that evidence
  3. Selects exactly one case based on semantic matching
  4. Exposes the corresponding action (continue, exit, or run Claude Code)

Core Principle

The LLM chooses the case. The YAML chooses the action.

The LLM is never trusted to invent actions—it only decides which case matches. The actual behavior always comes from your YAML.

Why You'd Use This

✅ Reduce Manual Triage

Instead of waiting for humans to debug CI failures, Decision Task can:

  • Auto-fix common issues (typos, formatting) via Claude Code
  • Block genuinely risky changes
  • Fast-track safe, low-risk PRs

✅ Understand Why, Not Just What

# Instead of:
if: ${{ job.status == 'failure' }}  # We know it failed, but not why

# You can ask:
cases:
  - name: trivial-typo
    if: Tests failed because of a one-line typo or syntax error.
    then: { run: { claudecode: { prompt: "Fix the typo." } } }
  
  - name: real-regression
    if: Tests failed because of a meaningful behavior change.
    then: { run: { exit: { code: 1, message: "Human review needed." } } }

✅ Audit Trail

Every decision is logged as JSON + Markdown:

  • What evidence was considered?
  • Which case was selected?
  • Why was the case selected? (model-provided explanation)
  • What action ran?

✅ Multi-Provider Support

Use Anthropic Claude, OpenAI, or offline mode (no API calls, pattern-based).

Quick Start (5 minutes)

1. Add the Action to Your Workflow

name: My CI Pipeline

on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm test > test.log 2>&1 || true  # Run tests, capture output
      
      - name: Decide how to handle failure
        id: decision
        uses: gianfa/decision-task@v0.0.1
        with:
          decision_file: .decisions/test-failure.yml
          llm_provider: anthropic
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
      
      - name: Auto-fix if needed
        if: steps.decision.outputs.action_type == 'claudecode'
        uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          prompt: ${{ steps.decision.outputs.claudecode_prompt }}
          allowed_tools: ${{ steps.decision.outputs.claudecode_allowed_tools }}
          max_turns: ${{ steps.decision.outputs.claudecode_max_turns }}

2. Create a Decision File (.decisions/test-failure.yml)

decision:
  name: test-failure-triage
  description: Decide how to handle test failures.

  sources:
    - "test.log"
    - "src/**/*.js"
    - "tests/**/*.js"

  on_no_match: block
  on_multiple_matches: block

  cases:
    - name: trivial-typo
      if: |
        Tests failed because of an obvious typo, syntax error, or one-line fix.
        Examples: missing semicolon, variable name misspelled, import statement wrong.
      then:
        run:
          claudecode:
            prompt: Fix the typo or syntax error causing test failures.
            allowed_tools: [Edit, Bash(npm test)]
            max_turns: 2

    - name: real-regression
      if: |
        Tests failed because the implementation behavior changed.
        The code logic is wrong, not just a typo.
      then:
        run:
          exit:
            code: 1
            message: Real regression detected. Manual review required.

    - name: flaky-test
      if: |
        Tests failed due to race conditions, timeouts, or environmental issues.
        The code is correct; the test execution is unreliable.
      then:
        run:
          continue: {}

3. Commit & Push

git add .decisions/test-failure.yml .github/workflows/...
git commit -m "feat: add test-failure decision"
git push

That's it! On the next push, Decision Task will evaluate your cases and route the workflow.

Real-World Example

See the complete ML risk quality gate example in examples/ml-quality-gate/:

  • Full working project with GitHub Actions workflow
  • Decision file that evaluates code risk
  • Demonstrates the full decision → action loop

Features

Feature Details
Multi-provider Use Anthropic Claude, OpenAI, or offline mode
Audit trail Every decision is logged as JSON + Markdown
Bounded evidence Configurable size limit (default: 40,000 characters)
Action types continue (proceed), exit (block), claudecode (auto-fix)
Fallback rules Define what happens if no case matches or multiple match
Offline mode No API keys required; uses pattern matching instead

How It Works

GitHub Event (push, PR, workflow_dispatch, etc.)
    ↓
[Decision Task Action]
    ├─ Read decision file (YAML)
    ├─ Collect evidence (files, logs, git context)
    ├─ Send to LLM: "Which case matches this evidence?"
    └─ LLM responds with: case name + reasoning
    ↓
[Workflow continues with the selected action]
    ├─ continue: proceed to next step
    ├─ exit: fail the workflow
    └─ claudecode: trigger Claude Code for auto-fixes

Documentation

FAQ

Q: What if the LLM gets it wrong?
A: You control the actions. Even if the LLM misclassifies a case, the only things that can happen are what you declared in your YAML. The LLM is never trusted to invent actions.

Q: Do I need an API key?
A: By default, yes (Anthropic or OpenAI). But you can use llm_provider: none for offline mode, where Decision Task matches a case named default if one exists.

Q: Can I auto-fix code with this?
A: Yes, via the claudecode action type. Route specific cases to Claude Code for automated edits (with your permission constraints).

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for setup instructions, commit conventions, and the release process.

License

Apache License 2.0 — See LICENSE for details.


Made with ❤️ for teams who want smarter CI/CD pipelines.

About

A lightweight GitHub Action for auditable, evidence-based LLM decisions in CI/CD pipelines.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages