Skip to content

Latest commit

 

History

History
177 lines (135 loc) · 10.6 KB

File metadata and controls

177 lines (135 loc) · 10.6 KB

Product Engineering Challenge Submission

Candidate


Run the project

Prerequisites

  • Python 3.10+ (Tested on Python 3.10, 3.11, 3.12, 3.13, 3.14)
  • No external libraries, build tools, or paid API keys required. The solution runs entirely offline using standard library capabilities (dataclasses, enum, typing, json, re, unittest).

Exact Run Commands

# 1. AC2, AC3, AC6: Multi-Step Root Cause Investigation (Status -> Logs -> Metrics -> Runbook -> Grounded Report)
python main.py --scenario multi-step

# 2. AC1: Single-Step Tool Selection & Evidence Grounding
python main.py --scenario single-step

# 3. AC4: Tool Failure & Adaptive Fallback (Simulated probe failure -> local log fallback)
python main.py --scenario tool-failure

# 4. AC5: Configured Step Limit Enforcement (Strict termination with MAX_STEPS_EXCEEDED)
python main.py --scenario step-limit

# 5. Output raw JSON trace summary for programmatic inspection
python main.py --scenario multi-step --json

# 6. Interactive Mode
python main.py --scenario interactive

Run the tests

Run the complete hermetic automated test suite (all 12 unit and integration tests passing in < 0.01s):

python -m unittest discover -s tests -p "test_*.py" -v

Architecture and data flow

The system decouples Orchestration / Control Loop, Tool Registry & Validation, Model Intelligence / Adapters, and Trace & Audit Logging:

                         ┌─────────────────────────────────────────┐
                         │               AgentLoop                 │
                         │  (Bounded step execution, state machine)│
                         └─────┬──────────────┬──────────────┬─────┘
                               │              │              │
                     ModelCalls│     ToolCalls│     Recording│
                               ▼              ▼              ▼
                      ┌─────────────┐  ┌─────────────┐  ┌─────────────┐
                      │ ModelAdapter│  │ ToolRegistry│  │ Execution   │
                      │ (Mock/LLM)  │  │ & Tools     │  │ Tracer      │
                      └─────────────┘  └─────────────┘  └─────────────┘
                                              │
                                              ▼
                                      ┌───────────────┐
                                      │  Type/Schema  │
                                      │  Validation   │
                                      └───────────────┘

Component Breakdown

  1. AgentLoop (src/agent.py):
    • Central state coordinator.
    • Enforces execution safety limits (max_steps_allowed).
    • Dispatches model decisions and executes tool calls.
    • Preserves state across iterations and catches tool/adapter runtime failures.
    • Sanitizes operational traces by redacting API tokens and sensitive credentials.
  2. Tool & ToolRegistry (src/tools.py):
    • Strongly-typed parameter schemas (ToolParameterSchema) with automated type, required-field, and enum validation.
    • Isolated concrete tools: SearchLogsTool, LookupMetricsTool, GetServiceStatusTool, QueryRunbookTool, and FailingToolMock.
    • Returns structured ToolResult with status (SUCCESS / ERROR), execution timing (duration_ms), and call IDs.
  3. ModelAdapter (src/model.py):
    • Pluggable interface (BaseModelAdapter).
    • IncidentInvestigatorDynamicAdapter: Heuristic multi-step investigator correlating service status, error logs, resource metrics, and runbooks.
    • ScriptedDeterministicAdapter: Exact, reproducible sequence executor for zero-flakiness automated testing.
  4. ExecutionTrace & StructuredAgentResponse (src/types.py):
    • Captures immutable, chronological records of every step (model rationale, tool inputs, outputs, errors, timings).
    • AC6 Decoupling: Strictly separates Grounded Tool Evidence (facts) from Root Cause Inferences and Remediation Recommendations.

Technology choices

  • Language: Python 3:
    • Chosen for direct alignment with AI agent harnesses and orchestration frameworks.
    • Rich standard library (dataclasses, typing, enum, abc.ABC, json, re) allows building a production-grade control loop without bloat or dependency drift.
  • Hermetic Testing with Zero External Dependencies:
    • Eliminates reliance on paid model APIs (OpenAI/Anthropic) for testing. Tests are fast, deterministic, and 100% offline.
  • Alternatives Considered:
    • TypeScript / Node.js: Excellent for web backends, but introduces package manager configs and build steps. Python was selected for maximum portability and native AI ergonomics.
    • Frameworks (LangChain / CrewAI / LlamaIndex): Deliberately avoided. Using heavy third-party agent frameworks obscures core control flow, increases overhead, and prevents fine-grained failure handling.

Important decisions

  1. Strict Decoupling of Grounded Tool Evidence vs. Model Inference (AC6):
    • In production incident triage, hallucinated "facts" cause severe outages. The StructuredAgentResponse data model enforces a strict separation between concrete tool observations (EvidenceItem) and agent hypotheses (root_cause_analysis / recommendations).
  2. Parameter Schema Validation at the Tool Boundary:
    • Rather than executing tools with unvalidated model arguments, Tool.validate_arguments() performs runtime type checking, enum verification, and default assignment. Schema errors are caught instantly and returned to the model as actionable feedback.
  3. Bounded Control Loop with Explicit Terminal States (MAX_STEPS_EXCEEDED):
    • The loop strictly enforces a maximum step count (max_steps). When exhausted, it halts immediately without issuing further model or tool calls, preventing infinite recursive loops and runaway operational costs.
  4. Zero-Crash Resilience & Secret Sanitization:
    • Tool exceptions (network timeouts, missing services, schema mismatches) never crash the orchestrator; they are encapsulated into ToolStatus.ERROR results so the agent can adapt or gracefully fail.
    • Automated regex scrubbing prevents bearer tokens, passwords, and API keys from leaking into execution traces.

Assumptions and limitations

  • Mocked Telemetry: The service tools currently query localized synthetic fixtures (src/data/synthetic_incident_data.py) representing a real-world payment service outage.
  • Synchronous Step Execution: Steps are executed sequentially. Parallel tool dispatching is supported by the ToolRegistry design but left for multi-threaded production harnesses.
  • Single Agent Context: Designed as a focused incident investigator loop rather than a complex hierarchical multi-agent supervisor.

Production and scale considerations

1. What prevents the agent from calling tools indefinitely?

  • A hard step ceiling (max_steps) enforced by the AgentLoop.
  • Cost/token budget tracking per session.
  • Cycle detection: The orchestrator can hash tool signatures and arguments to identify identical repeated calls (looping traps).

2. How would you add a consequential tool requiring human approval?

  • Introduce an approval_required: bool flag on ToolParameterSchema or ToolDefinition (e.g. for RollbackDeploymentTool or RestartClusterTool).
  • When the model selects a high-consequence tool, the loop transitions to LoopStatus.AWAITING_HUMAN_APPROVAL, emits an event to an operational queue (e.g. Slack/PagerDuty/Web UI), and suspends execution state.
  • Upon human confirmation or rejection, the loop resumes with the decision injected as a tool outcome.

3. How would you run concurrent agent jobs in a cloud environment?

  • State Externalization: Store ExecutionTrace and agent checkpoint state in a distributed database (PostgreSQL / DynamoDB) or Redis stream.
  • Worker Execution: Deploy worker containers using Temporal, Celery, or AWS SQS.
  • Idempotency Keys: Use session_id and step_number as idempotency keys to ensure at-most-once tool execution in case of worker restarts.

4. Which run data would you persist for debugging, cost analysis, and evaluation?

  • Operational Traces: Complete ExecutionTrace JSON records with step timings, tool inputs, and error states.
  • Token & Cost Metrics: Prompt tokens, completion tokens, latency per step, and cost estimates.
  • Golden Evaluation Datasets: Run historical incident scenarios against the harness continuously in CI/CD to measure tool selection accuracy and hallucination rates.

AI usage

  • AI Tools Used: Google DeepMind Antigravity / Gemini.
  • How they contributed:
    • Used for rapid scaffolding of synthetic incident fixtures, domain type boilerplate, and initial unit test templates.
    • All architecture decisions, validation contracts, bounded control loop design, and trade-off analyses were engineered and verified.

Credibility note

High-Throughput Fintech Event Ingestion & Incident Telemetry Engine

  • Problem Solved: Built and shipped a high-concurrency payment webhook and event dispatch engine handling financial transaction updates across multiple third-party payment gateways.
  • Personal Contribution:
    • Designed the core idempotent message ingestion pipeline and exponential backoff retry scheduler.
    • Implemented dead-letter queues (DLQ), circuit breakers, and telemetry collectors for real-time error rate tracking.
  • Scale / Operational Complexity:
    • Handled sustained loads of 2,500+ events/sec with p99 delivery latency under 80ms.
    • Maintained zero data loss during upstream third-party gateway downtime by queuing and replaying millions of backlogged webhook events with rate-limiting.
  • Difficult Engineering Decision:
    • Chose distributed Redis-backed sliding-window rate limiters combined with transactional outbox tables in PostgreSQL instead of relying solely on in-memory queues, ensuring delivery guarantees survived process crashes and pod restarts.