Sentinel is a full-stack, distributed API uptime and latency monitoring platform built to demonstrate production-style backend and distributed systems concepts.
Users can register HTTP/HTTPS endpoints, monitor them from multiple regions, inspect historical checks and latency metrics, track automatically detected incidents, receive real-time dashboard updates, and optionally request AI-assisted health and incident analysis.
Sentinel keeps monitoring and incident decisions deterministic. AI runs asynchronously and never decides whether a monitor is healthy, degraded, or down.
- Multi-region HTTP/HTTPS API monitoring
- Distributed probe workers
- BullMQ + Redis background job processing
- Automated scheduling of API checks
- MongoDB persistence
- Automatic incident detection and recovery
- Real-time dashboard updates using Socket.IO
- JWT authentication
- User-specific monitor isolation
- API latency tracking
- Uptime calculation
- p50, p95, and p99 latency metrics
- Configurable failure and recovery thresholds
- SSRF-aware outbound HTTP requests
- Optional AI-powered health insights
- Optional AI-generated incident summaries
- React dashboard
- Unit and integration tests with Vitest
Sentinel uses a modular monolith with multiple independent runtime processes.
The backend shares a single TypeScript codebase, while different processes handle API requests, scheduling, monitoring, incident detection, and AI analysis.
┌──────────────────────┐
│ React Dashboard │
└──────────┬───────────┘
│
│ HTTP / Socket.IO
▼
┌──────────────────────┐
│ Express API │
│ + Socket.IO │
└───────┬───────┬──────┘
│ │
MongoDB │ │ Redis
│ │
▼ ▼
┌───────────┐ ┌─────────────┐
│ MongoDB │ │ Redis │
│ Database │ │ + BullMQ │
└───────────┘ └──────┬──────┘
│
┌────────────────────┼─────────────────────┐
│ │ │
▼ ▼ ▼
┌────────────────┐ ┌────────────────┐ ┌────────────────┐
│ Mumbai Worker │ │Singapore Worker│ │Frankfurt Worker│
└───────┬────────┘ └───────┬────────┘ └───────┬────────┘
│ │ │
└────────────────────┼────────────────────┘
│
▼
Monitored APIs
Redis / BullMQ
│
┌───────────────┴────────────────┐
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Incident Worker │ │ AI Worker │
└─────────────────┘ └─────────────────┘
| Process | Responsibility |
|---|---|
| API Server | REST API, authentication, monitor CRUD, metrics, incidents and Socket.IO |
| Scheduler | Finds monitors that need to be checked and creates regional probe jobs |
| Probe Worker | Performs HTTP checks from a configured region |
| Incident Worker | Determines health state and creates/resolves incidents |
| AI Worker | Handles optional AI analysis requests |
| Frontend | Provides the monitoring dashboard and authentication UI |
- Node.js
- TypeScript
- Express
- MongoDB
- Mongoose
- Redis
- BullMQ
- Socket.IO
- Zod
- Pino
- JWT
- Argon2
- OpenAI SDK
- Vitest
- Supertest
- React
- TypeScript
- Vite
- React Router
- TanStack Query
- React Hook Form
- Zod
- Tailwind CSS
- Motion
- Socket.IO Client
- Vitest
- Testing Library
Sentinel/
│
├── backend/
│ ├── src/
│ │ ├── api/
│ │ │ └── Express server, middleware and Socket.IO
│ │ │
│ │ ├── config/
│ │ │ └── Environment configuration and logging
│ │ │
│ │ ├── database/
│ │ │ └── MongoDB connection and models
│ │ │
│ │ ├── modules/
│ │ │ ├── authentication
│ │ │ ├── monitors
│ │ │ ├── metrics
│ │ │ ├── incidents
│ │ │ └── AI
│ │ │
│ │ ├── monitoring/
│ │ │ ├── HTTP checker
│ │ │ ├── transport
│ │ │ └── SSRF protection
│ │ │
│ │ ├── queues/
│ │ │ └── BullMQ queues and job contracts
│ │ │
│ │ ├── realtime/
│ │ │ └── Redis and Socket.IO events
│ │ │
│ │ ├── scheduler/
│ │ │ └── Monitor scheduling loop
│ │ │
│ │ └── workers/
│ │ ├── probe worker
│ │ ├── incident worker
│ │ └── AI worker
│ │
│ └── tests/
│
├── frontend/
│ └── src/
│ ├── api/
│ ├── app/
│ ├── auth/
│ ├── components/
│ ├── features/
│ ├── pages/
│ └── realtime/
│
└── README.md
Before running Sentinel locally, make sure you have:
- Node.js 20+
- npm
- MongoDB
- Redis
By default, Sentinel expects:
MongoDB: mongodb://localhost:27017
Redis: redis://localhost:6379
Clone the repository:
git clone <your-repository-url>
cd Sentinelcd backend
npm installcd ../frontend
npm installNavigate to:
cd backendCopy the example environment file:
cp .env.example .envExample configuration:
NODE_ENV=development
PORT=4000
LOG_LEVEL=debug
MONGODB_URI=mongodb://localhost:27017/sentinel
REDIS_URL=redis://localhost:6379
BULLMQ_PREFIX=sentinel
JWT_SECRET=replace-with-a-long-random-secret
JWT_EXPIRES_IN=7d
CLIENT_ORIGIN=http://localhost:5173
ENABLED_REGIONS=mumbai,singapore,frankfurt
PROBE_REGION=mumbai
PROBE_CONCURRENCY=20
SCHEDULER_POLL_INTERVAL_MS=5000
GLOBAL_CHECK_TIMEOUT_MS=10000
MAX_RESPONSE_BODY_BYTES=65536
ALLOW_PRIVATE_NETWORK_TARGETS=false
AI_ENABLED=false
OPENAI_API_KEY=your-openai-api-key
OPENAI_MODEL=your-supported-model
AI_REQUEST_TIMEOUT_MS=30000Use a strong random value for:
JWT_SECRETAI configuration is only required when:
AI_ENABLED=trueNavigate to:
cd frontendCopy the environment example:
cp .env.example .envExample:
VITE_BACKEND_URL=http://localhost:4000Sentinel intentionally runs several backend processes independently.
For local development, open separate terminals for each process.
Make sure MongoDB is running locally.
The default connection is:
mongodb://localhost:27017/sentinel
Make sure Redis is running on:
redis://localhost:6379
cd backend
npm run devThe backend will be available at:
http://localhost:4000
API base URL:
http://localhost:4000/api/v1
Health endpoint:
GET http://localhost:4000/api/v1/health
Open another terminal:
cd backend
npm run dev:schedulerThe scheduler periodically looks for monitors that need to be checked.
Sentinel can run multiple regional workers.
cd backend
PROBE_REGION=mumbai npm run dev:probecd backend
PROBE_REGION=singapore npm run dev:probecd backend
PROBE_REGION=frankfurt npm run dev:probeOn PowerShell:
$env:PROBE_REGION="mumbai"
npm run dev:probeEach worker consumes only jobs assigned to its configured region.
In local development, all workers can run on the same machine.
In production, they can be deployed to different geographic locations.
cd backend
npm run dev:incidentThe incident worker analyzes recent monitoring results and determines whether a monitor should be:
healthy
degraded
down
It also automatically creates and resolves incidents.
The AI worker is optional.
First enable AI:
AI_ENABLED=trueThen configure:
OPENAI_API_KEY=your-key
OPENAI_MODEL=your-modelStart the worker:
cd backend
npm run dev:aiIf AI is disabled, the monitoring system continues to work normally.
cd frontend
npm run devOpen:
http://localhost:5173
With the default three monitoring regions, a complete development environment consists of:
MongoDB
Redis
API Server
Scheduler
Mumbai Probe Worker
Singapore Probe Worker
Frankfurt Probe Worker
Incident Worker
Frontend
The AI worker can optionally be added.
When a user creates a monitor, Sentinel follows this flow:
Create Monitor
│
▼
Scheduler detects due monitor
│
▼
Creates one probe job per region
│
▼
Regional probe workers execute HTTP checks
│
▼
Check results stored in MongoDB
│
▼
Incident evaluation job created
│
▼
Incident worker evaluates recent results
│
▼
Monitor status updated
│
▼
Realtime event published
│
▼
React dashboard updates
In more detail:
- A user creates a monitor.
- The monitor contains a URL, interval, timeout, regions and thresholds.
- The scheduler finds monitors whose next check time has arrived.
- A BullMQ job is generated for each configured region.
- Regional probe workers perform the HTTP request.
- Results are stored in MongoDB.
- An incident evaluation job is queued.
- The incident worker examines recent regional results.
- Monitor health is recalculated.
- Incidents are opened or resolved if required.
- Redis publishes a real-time domain event.
- Socket.IO sends the update to the authenticated dashboard.
- The frontend refreshes the appropriate data.
Each monitor can contain:
- Monitor name
- HTTP/HTTPS URL
- HTTP method
- Check interval
- Request timeout
- Expected HTTP status codes
- Latency warning threshold
- Failure threshold
- Recovery threshold
- Enabled regions
- Active or paused state
Supported methods include:
GET
HEAD
The minimum monitoring interval is currently:
10 seconds
The maximum per-monitor timeout is:
30 seconds
A monitor can have the following states:
pending
healthy
degraded
down
paused
Example lifecycle:
pending
│
▼
healthy
│
├─────────────► degraded
│ │
│ ▼
│ down
│ │
└──────────────────┘
Paused monitors use:
paused
Incident detection is intentionally deterministic.
Sentinel does not use AI to decide whether an API is down.
A single failed HTTP request does not necessarily create an incident.
Instead, Sentinel evaluates:
- Recent checks
- Regional results
- Consecutive failures
- Consecutive successful checks
- Configured failure threshold
- Configured recovery threshold
This helps avoid creating incidents for temporary network failures.
The incident worker also prevents duplicate active incidents.
Once the recovery criteria are satisfied, the existing incident is automatically resolved.
Sentinel calculates monitoring statistics such as:
- Uptime percentage
- Average latency
- p50 latency
- p95 latency
- p99 latency
- Recent check history
- Regional health
- Recent incidents
Example:
Uptime: 99.95%
Avg Latency: 182 ms
p50: 160 ms
p95: 310 ms
p99: 470 ms
All metrics are scoped to the authenticated monitor owner.
All API endpoints are available under:
/api/v1
POST /api/v1/auth/registerCreates a new user account.
POST /api/v1/auth/loginAuthenticates the user and returns a JWT.
GET /api/v1/auth/meReturns the currently authenticated user.
POST /api/v1/monitorsGET /api/v1/monitorsGET /api/v1/monitors/:monitorIdPATCH /api/v1/monitors/:monitorIdDELETE /api/v1/monitors/:monitorIdPOST /api/v1/monitors/:monitorId/pausePOST /api/v1/monitors/:monitorId/resumeGET /api/v1/monitors/:monitorId/checksReturns paginated monitoring history.
GET /api/v1/monitors/:monitorId/metricsReturns uptime and latency statistics.
GET /api/v1/monitors/:monitorId/incidentsGET /api/v1/incidents/:incidentIdReturns the incident and its timeline.
POST /api/v1/monitors/:monitorId/ai-insightsQueues an asynchronous AI health analysis.
POST /api/v1/incidents/:incidentId/ai-summaryQueues an AI-generated incident summary.
GET /api/v1/ai-analyses/:analysisIdReturns the current analysis state and result.
Sentinel uses Socket.IO for live dashboard updates.
Supported events include:
check.completed
monitor.status_changed
incident.opened
incident.resolved
ai.analysis.completed
ai.analysis.failed
Workers publish domain events through Redis.
The API server receives those events and forwards them to authenticated Socket.IO clients.
This allows backend workers to remain independent from the API server's process memory.
Sentinel contains optional AI functionality.
AI is deliberately isolated from the monitoring system's critical decision path.
Users can request an AI-generated summary of recent monitor health and latency information.
The model receives bounded monitoring telemetry rather than unrestricted raw data.
Users can request an AI-generated explanation of a resolved incident using:
- Incident timeline
- Regional checks
- Latency information
- Failure patterns
- Recovery information
AI functionality follows several rules:
- AI is explicitly requested by the user.
- AI runs asynchronously.
- AI does not determine monitor health.
- AI does not open incidents.
- AI does not resolve incidents.
- Provider failures do not break monitoring.
- AI output is validated before persistence.
- Response bodies are not unnecessarily sent to the AI provider.
- Monitoring continues normally when AI is disabled.
Sentinel contains several protections that are especially important for an uptime monitoring platform.
Protected endpoints require JWT authentication.
Passwords are hashed using:
Argon2
Users can only access their own:
- Monitors
- Checks
- Metrics
- Incidents
- AI analyses
Socket.IO connections are also authenticated and user-scoped.
Monitoring platforms make HTTP requests to user-provided URLs, which can create SSRF risks.
Sentinel protects against this by validating monitoring destinations.
By default:
ALLOW_PRIVATE_NETWORK_TARGETS=falsePrivate network destinations are blocked.
Monitor URLs:
- Must use HTTP or HTTPS
- Cannot contain embedded authentication credentials
- Are checked before requests are performed
- Have redirect destinations validated
HTTP checks are bounded by:
GLOBAL_CHECK_TIMEOUT_MS
MAX_RESPONSE_BODY_BYTESThis prevents monitored endpoints from causing unlimited response reads or indefinitely hanging worker requests.
cd backend
npm testnpm run test:integrationnpm run typechecknpm run lintcd frontend
npm testnpm run typechecknpm run lintBuild the backend:
cd backend
npm run buildStart the main API:
npm startOther production runtime commands include:
npm run start:schedulernpm run start:probenpm run start:incidentnpm run start:aiEach probe worker still requires its own:
PROBE_REGIONFor example:
PROBE_REGION=mumbai npm run start:probeBuild the frontend:
cd frontend
npm run buildVite generates the production frontend inside:
frontend/dist/
Sentinel focuses on the core engineering challenges involved in building a distributed monitoring platform.
The current version intentionally does not attempt to reproduce every feature of commercial monitoring systems.
Possible future additions include:
- Email notifications
- Slack notifications
- Discord notifications
- Webhook alerts
- Public status pages
- Team accounts
- Role-based access control
- Browser synthetic monitoring
- TCP monitoring
- ICMP monitoring
- SSL certificate expiration monitoring
- Domain expiration monitoring
- Kubernetes deployment
- Billing and subscription plans
- Additional monitoring regions
- Alert escalation policies
- Scheduled maintenance windows
Sentinel was built to explore engineering concepts that appear in real production systems.
The project demonstrates:
- Distributed background processing
- Queue-based architecture
- Worker concurrency
- Multi-region processing
- Asynchronous workloads
- Idempotency
- Failure isolation
- Real-time systems
- State machines
- Incident lifecycle management
- API authentication
- User resource isolation
- MongoDB data modeling
- Operational metrics
- SSRF protection
- Safe outbound HTTP handling
- AI integration outside critical infrastructure decisions
Rather than building only a CRUD dashboard, Sentinel focuses on the infrastructure and failure-handling problems involved in creating a real monitoring service.
This project is intended primarily for educational and portfolio purposes.
If you found this project useful, consider giving the repository a ⭐.