Complete open-source monitoring and observability platform. https://oneuptime.com
  • TypeScript 85.8%
  • EJS 7%
  • JavaScript 5.1%
  • Shell 0.8%
  • Go 0.4%
  • Other 0.9%
Find a file
2026-08-05 15:41:10 +01:00
.github Merge branch 'claude/angry-saha-fca840': keep the agent's monitor secret key out of its world-readable log file 2026-08-05 12:59:08 +01:00
.husky feat: Enhance alert and incident label rule engines to inherit labels from Docker hosts and Kubernetes clusters 2026-05-25 14:19:51 +01:00
.vscode Remove IsolatedVM service and related configurations from the project 2026-03-03 12:25:31 +00:00
App Merge remote-tracking branch 'origin/master' into claude/user-table-duplicates-9fb6ec 2026-08-05 14:53:55 +01:00
Backups Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
CephAgent Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
Certs Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
CLI chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
Clickhouse feat(clickhouse): add distributed insert tuning configuration to optimize telemetry ingestion 2026-07-01 08:20:57 -07:00
Common Merge remote-tracking branch 'origin/master' into claude/user-table-duplicates-9fb6ec 2026-08-05 14:53:55 +01:00
Data Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
Devops Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
DockerAgent Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
DockerSwarmAgent fix(swarm): repair the Docker Swarm agent install 2026-07-21 15:01:20 +01:00
E2E chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
Environment Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
Examples chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
FluentBit Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
Fluentd Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
HelmChart Merge pull request #3001 from arvid-blikk/feat/helm-teams-existing-secret 2026-08-05 13:39:47 +01:00
Home chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
InfrastructureAgent docs(agent): correct stale comments claiming the secret-key leak is unfixed 2026-08-05 15:03:58 +01:00
Internal/Roadmap docs(roadmap): remove the AI Sentinel vision and execution docs 2026-08-04 11:14:48 +01:00
KubernetesCostAgent chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
KubernetesLogTailer chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
MobileApp chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
Nginx chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
OTelCollector fix: update nodemailer and its types to latest versions across packages 2026-06-12 13:16:50 +03:00
PodmanAgent add podman agent 2026-06-15 11:41:10 +03:00
Probe chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
ProxmoxAgent feat: Proxmox and Ceph world-class pass (v3) 2026-06-13 01:41:50 +03:00
Runner chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
Scripts chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
SslCertificates
Tests fix: update Node.js version to 26 across all workflows and Dockerfiles for improved compatibility and security 2026-06-12 14:49:01 +03:00
TestServer chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
translations docs: collapse the 'automate the busywork' workflows screenshot into a disclosure 2026-07-08 14:11:27 +01:00
.bash_profile
.dockerignore feat(dockerignore): add large directories to ignore for Docker builds 2026-01-22 21:42:52 +00:00
.gitattributes Add bulk status page monitor selection 2026-08-03 17:08:02 +01:00
.gitignore fix(ui): load Monaco from the install instead of the CDN 2026-07-27 19:49:34 +05:30
.prettierignore refactor: remove APIReference from nodemon watch and docker-compose volumes 2026-02-22 13:46:44 +00:00
AGENTS.md ci(postgres): fail CI when the schema drifts from the entities 2026-08-04 22:30:45 +01:00
babel.config.ts
backup.sh
CHANGELOG
CLAUDE.md Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
clean-npm-install.sh
code-of-conduct.md Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
config.example.env feat(runner): merge the Runbook Agent and AI Agent into the OneUptime Runner 2026-08-03 18:22:07 +01:00
configure.sh refactor: update environment variable loading in installation scripts 2026-03-06 13:34:35 +00:00
CONTRIBUTING.md Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
docker-compose.base.yml feat(runner): merge the Runbook Agent and AI Agent into the OneUptime Runner 2026-08-03 18:22:07 +01:00
docker-compose.billing.yml Refactor code structure for improved readability and maintainability 2026-06-22 22:21:35 +01:00
docker-compose.dev.yml feat(runner): merge the Runbook Agent and AI Agent into the OneUptime Runner 2026-08-03 18:22:07 +01:00
docker-compose.e2e.yml feat: add E2E testing support with docker-compose configuration 2025-10-09 11:39:30 +01:00
docker-compose.yml fix(compose): rename stale ai-agent service to runner in docker-compose.yml 2026-08-04 02:13:44 +01:00
eslint.config.js chore(lint): ignore .claude/worktrees 2026-07-17 09:00:28 +01:00
install-node-modules.sh
install.sh refactor: update environment variable loading in installation scripts 2026-03-06 13:34:35 +00:00
LICENSE
MAINTAINERS
migration-create.sh
migration-run.sh
npm-audit-fix.sh chore(ci): don't mark whole run as failed when npm audit fix errors; only report the error 2025-10-29 16:26:40 +00:00
package-lock.json chore: npm audit fix 2026-07-09 03:32:44 +00:00
package.json chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00
README.md feat(ai): sequence remediation after RCA, auto-open verified fix PRs 2026-08-04 20:33:56 +01:00
remove-node-modules.sh
restore.sh
SECURITY.md
tsconfig.json
uninstall.sh
update-node-modules.sh
update.sh
VERSION chore(version): bump version to 12.0.1 2026-08-05 08:20:07 +01:00

English · 简体中文 · 繁體中文 · 日本語 · 한국어 · Español · Français · Deutsch · Português · Italiano · Русский · हिन्दी · Nederlands · Dansk · Svenska · Norsk

OneUptime logo

Agentic observability — one open-source platform for uptime, incidents, on-call, status pages, logs, traces, metrics & APM.

When things go wrong, be the first to know — and the fastest to fix.

OneUptime replaces a whole shelf of SaaS tools with one platform you can self-host for free. It catches the outage, pages the right person, updates your status page, finds the root cause, and even opens the fix PR.

License Release Stars Helm Chart Slack

Website  •  Docs  •  Quick Start  •  Pricing  •  Contribute

🚀 Try OneUptime Cloud — free forever plan, no credit card →


OneUptime command center during a live incident

Replace your whole observability stack

OneUptime brings monitoring, alerting, incident response, and observability into a single open-source app — so you stop paying for (and stitching together) a dozen separate tools.

Instead of… Use OneUptime for…
Pingdom / UptimeRobot Uptime Monitoring — website, API, ping, port, SSL, DNS & synthetic checks from around the world
StatusPage.io Status Pages — branded public & private status pages with subscribers
PagerDuty / Opsgenie On-Call & Alerts — schedules, escalation policies, SMS / call / push / Slack
Incident.io Incident Management — declare, triage, communicate, and post-mortem
Datadog / New Relic APM & Metrics — traces, dashboards, and service performance
Loggly Log Management — collect, search, and alert on logs
Sentry Error Tracking — exceptions with full stack traces and context

All of it is 100% open source (Apache 2.0) and free to self-host.


🌙 One incident, handled end to end

It's 2:47 AM. Checkout starts timing out. Here's what OneUptime does before most tools would even fire the first alert — and what the screenshots below actually show.

1 · Detect — know in seconds

Probes in multiple regions catch checkout latency blowing past your 5s threshold and open an incident automatically — before your customers hit refresh.

Detect — global monitoring catches the checkout API degrading

2 · Respond — the right person, paged

The on-call engineer for the Payments policy is called, texted, and push-notified, escalating to backup automatically until someone acknowledges.

Respond — the incident is routed to on-call and acknowledged

3 · Communicate — customers in the loop

Your status page updates itself and every subscriber is notified by email and SMS — no one has to hand-write the update.

Communicate — the public status page updates and notifies subscribers

4 · Diagnose — root cause, found

Traces, logs, and metrics are correlated down to the exact span: a slow SELECT … FOR UPDATE on orders, stuck on a missing index.

Diagnose — the trace waterfall pinpoints the slow database span

5 · Auto-Fix — the fix, drafted for you

The AI agent opens a pull request with the fix, linked to the incident, verified against your repository's configured build and test commands before it opens — you review and merge. Like an SRE that never sleeps.

Auto-Fix — the AI agent opens a pull request with the fix


Quick Start

☁️ OneUptime Cloud — the easy way

Zero setup, always up to date, and it funds the open-source project.

Sign up free at oneuptime.com

🐳 Self-host with Docker Compose

Everything you need on a single server (Debian / Ubuntu / RHEL, Docker + Docker Compose). Great for homelabs and small teams — a Raspberry Pi even works.

# 1. Clone the release branch
git clone --depth 1 --single-branch --branch release https://github.com/OneUptime/oneuptime.git
cd oneuptime

# 2. Create your config (then edit it — set strong, random secrets!)
cp config.example.env config.env

# 3. Start everything
npm start

OneUptime is now running at http://localhost — open it and create your first account.

📖 Full guide: Docker Compose install · Sizing & requirements

☸️ Kubernetes with Helm — for production

helm repo add oneuptime https://helm-chart.oneuptime.com
helm install oneuptime oneuptime/oneuptime

📖 Full install instructions & values on Artifact Hub →

Upgrading an existing install? See the upgrade guide.


Everything in the box

Feature What it does
📊 Uptime Monitoring Website, API, IP, port, SSL, DNS, and synthetic monitors from multiple global regions.
📋 Status Pages Beautiful branded status pages, incident history, scheduled maintenance, and subscriber notifications.
🚨 Incident Management End-to-end incident workflow: declare, assign, communicate, resolve, and run post-mortems.
📞 On-Call & Alerts On-call schedules and escalation policies with SMS, phone call, push, email, and Slack alerts.
📝 Log Management Ingest, store, search, and alert on logs via OpenTelemetry.
🔍 APM & Traces Distributed traces, spans, and performance dashboards to find slow paths and bottlenecks.
📈 Metrics & Dashboards Custom dashboards over your telemetry — build the views your team needs.
🐛 Error Tracking Capture exceptions with full stack traces, context, and release tracking.
Workflows Automate and integrate with Slack, Jira, GitHub, Microsoft Teams, and 5,000+ apps.
🤖 AI Copilot An always-on agent that finds anomalies across logs, traces & metrics, spots root causes, and opens PRs with fixes.
Automate the busywork

Wire up escalations, ticketing, and notifications on a visual, no-code canvas — or drop in custom code. The incident above paged on-call, opened a Jira ticket, and posted to Slack without anyone lifting a finger.

Workflows — a no-code automation canvas for incident escalation

🖥️ Infrastructure Monitoring

Drop in copy-paste, OpenTelemetry-based agents to watch everything your services run on — with ready-made alert templates included:

  • Servers & VMs — CPU, memory, disk, network, processes, and logs from Linux, macOS & Windows. Docs →
  • Kubernetes — one helm install ships node/pod/container/cluster metrics, events, logs, and eBPF traces & service maps. Docs →
  • Docker — a single agent auto-discovers every container and ships metrics & logs. Docs →
  • Podman — same one-agent auto-discovery via Podman's Docker-compatible socket. Docs →
  • Proxmox — nodes, VMs, containers, storage, HA state, backup coverage & replication health. Docs →
  • Ceph — cluster health, capacity forecasts, and OSD/pool/PG/monitor visibility. Docs →

💼 Community vs. Enterprise

Community Enterprise
Best for Self-hosters & small teams Regulated teams needing premium support
Cost Free & open source Contact sales
Features Full feature set Full feature set + hardened images, priority support, custom features & data residency

💡 Why OneUptime?

Our mission is simple: reduce downtime and help more products succeed. Instead of duct-taping seven vendors together, you get one platform that helps you understand why things break, respond to incidents fast, and cut operational toil — fully open source, so you own your data and your stack.


🤝 Contributing

We welcome contributions of every size. Start here:

❤️ Support the project

If OneUptime is useful to you:

  • Star this repo — it genuinely helps others find us
  • 💵 Sponsor us — every dollar ships new features
  • 🛍️ Grab some merch — all proceeds fund open-source development

📄 License

OneUptime is licensed under the Apache License 2.0.

Made with ❤️ by the OneUptime team and contributors.