Meet OpenUptime, an open source uptime monitor you can run yourself认识 OpenUptime:一个可以自己部署的开源可用性监控
What an open source uptime monitor does when your site goes down at 3 a.m., what it keeps an eye on, and how to run it: hosted, in Docker, or on your own Cloudflare account.凌晨 3 点网站挂了,一个开源的可用性监控会做什么、它盯着哪些东西,以及怎么用起来:托管版、Docker,或者部署在你自己的 Cloudflare 账号上。

It is 3:12 in the morning and your site is down. The only question that matters is who finds out first: you, or the customer who is already typing an angry message to your support inbox.凌晨 3 点 12 分,你的网站挂了。真正要紧的只有一个问题:是你先知道,还是已经在给客服写投诉的那位客户先知道。
Uptime monitoring exists so that it is always you. It is a boring job, and we like boring jobs: something pokes your service every so often, and when it stops answering, it taps you on the shoulder. OpenUptime is our take on that job, and it is open source.可用性监控存在的意义,就是让"先知道"的永远是你。这是个无聊的活儿,但我们喜欢无聊的活儿:隔一会儿戳一下你的服务,它不应声了,就来拍你一下肩膀。OpenUptime 就是我们对这件事的做法,并且它是开源的。
Why another one?为什么又做一个?
Honestly, because the choices felt lopsided. On one side there are hosted products you rent for as long as you care about uptime, which is to say forever. On the other side there is a cron job, a curl command and a Slack webhook you wrote on a Friday afternoon and have been quietly babysitting since.说实话,是因为现有的选择有点两极分化。一边是托管产品,你只要还在乎可用性就得一直租,也就是永远。另一边是某个周五下午你自己写的 cron、curl 加一个 Slack webhook,从那以后就一直偷偷照看着它。
We wanted the middle: a complete product with monitors, incidents, alerts and a status page, that you can run yourself when you want to, and that does not fall apart when you do.我们想要的是中间地带:监控、故障事件、告警、状态页一应俱全的完整产品,想自己跑的时候就能自己跑,而且自己跑起来也不会散架。
What it keeps an eye on它帮你盯什么
The usual suspect is a web page, so HTTP monitors are the workhorse. You can check the status code, and also that the page contains (or does not contain) a given word, which catches the classic "returns 200, shows an error page" situation. Custom request headers are there for the endpoints that want a bearer token.最常见的对象是网页,所以 HTTP 监控是主力。可以检查状态码,也可以检查页面里"有没有"或"不该有"某个词,专治"返回 200,页面却是一张报错页"这种经典情况。需要 Bearer 令牌的接口,可以加自定义请求头。
Beyond that:除此之外:
- TCP and DNS for the things that are not web pages but still need to be up.TCP 和 DNS:不是网页,但同样不能挂的东西。
- Postgres, MySQL and Redis reachability checks. They do the handshake (or a
PING) and never send credentials, so there is no password to leak.Postgres、MySQL、Redis 连通性检查:只做握手(或PING),不发送任何凭据,也就没有密码可泄露。 - Heartbeats for cron jobs and background workers. Instead of us calling you, your job calls us; if it goes quiet for longer than the grace period you set, that counts as a failure. Handy for the nightly backup that has been silently failing since March.心跳监控:给定时任务和后台进程用。不是我们去找你,而是你的任务来找我们;超过你设的宽限时间还没动静,就算失败。特别适合那个从三月起就在悄悄失败的夜间备份。
What happens when something breaks出事的时候会发生什么
One failed check is not an outage. Networks hiccup. So each monitor has its own "how many failures in a row before we call it down" setting, and an incident opens only after that. When the checks start passing again, the incident resolves by itself. Nobody has to remember to close a ticket at 4 a.m.一次检查失败不算故障,网络偶尔也会打个嗝。所以每个监控都有自己的"连续失败几次才算挂"的设置,达到之后才会创建故障事件;检查重新通过,事件就自动恢复。不会有人半夜 4 点还得记着去关工单。
There is also a middle state for the "it works, but it is crawling" days. Each monitor has a slow threshold (1,500 ms by default). Go over it and the monitor shows as degraded; you can optionally open a minor incident after a few slow checks in a row.还有一个中间状态,留给"能用,但慢得像在爬"的日子。每个监控都有一个慢响应阈值(默认 1500 毫秒),超过了就显示为"降级";你也可以设置连续几次变慢后,自动开一个轻微故障。
Alerts can go to email, Slack, Lark (Feishu), Telegram, PagerDuty or any webhook, and you can choose per monitor which channels it should bother. The database being down and the marketing page being slow probably should not wake the same people.告警可以发到邮件、Slack、飞书、Telegram、PagerDuty,或者任意 Webhook,每个监控还能单独选择通知哪些渠道。数据库挂了和营销页变慢,大概不该吵醒同一批人。
Planned maintenance? Create a maintenance window. Checks keep running so you still have the data, but nobody gets paged for something you did on purpose.有计划内的维护?建一个维护窗口就行。检查照常进行,数据还在,但不会为你自己故意做的事去吵醒任何人。
The status page, so people stop asking you状态页:让大家别再来问你
Half the pain of an outage is the "is it down?" messages. A status page takes most of that away. You pick which monitors to show, post updates as you work on the incident, and visitors can subscribe by email (with a proper double opt-in and a one-click unsubscribe) or follow an RSS feed. When an incident opens or resolves, subscribers hear about it without you writing a single email.一次故障里,一半的痛苦来自"是不是挂了?"这类消息。状态页能把大部分消息挡掉。你选择展示哪些监控,处理过程中发布更新,访客可以用邮箱订阅(双重确认,每封邮件都带一键退订),也可以订阅 RSS。事件创建或恢复时,订阅者自动收到通知,你一封邮件都不用写。
Your AI agent gets a seat too你的 AI 助手也有一个座位
OpenUptime ships an MCP server. Point Claude (or any MCP client) at it and you can say "which monitors are down?", "run that check now" or "post an update that we are investigating", and it just does it. You decide what the agent may touch: every workspace, selected ones, or only certain projects, and you can revoke it later. The MCP guide has the details.OpenUptime 自带 MCP 服务。把 Claude(或任何 MCP 客户端)接上去,你可以直接说"现在哪些监控挂了?""立刻跑一下那个检查""发一条我们正在排查的更新",它就会去做。agent 能碰什么由你决定:全部工作区、指定工作区,或者只限某几个项目,之后随时可以撤销。细节见 MCP 文档。
How it runs, in one breath它是怎么跑起来的,一口气说完
A Cloudflare Worker wakes up every minute, picks the monitors that are due, hands them out in packs of 25 through a queue, and another Worker invocation runs each pack, writes the results and sends alerts. Data lives in an ordinary Postgres. That last choice raised some eyebrows, so we wrote about it: why Postgres and not D1.一个 Cloudflare Worker 每分钟醒一次,挑出到点的监控,按每 25 个一组通过队列分发出去,另一个 Worker 调用负责执行每一组、写入结果、发送告警。数据放在普通的 Postgres 里。最后这个选择让一些人挑了眉毛,所以我们专门写了一篇:为什么用 Postgres 而不是 D1。
Hosted, or on your own box用托管版,或者放在自己的机器上
Hosted at openuptime.app: sign up and you are monitoring in a couple of minutes. The free plan gives you 3 monitors checked every 5 minutes, which is plenty for a side project. Paid plans add more monitors, one-minute checks and more notification email; see pricing.托管版在 openuptime.app:注册后几分钟就能开始监控。免费套餐包含 3 个监控、每 5 分钟检查一次,做个人项目够用了。付费套餐有更多监控、1 分钟检查和更多通知邮件,见价格页。
Self-hosted: OpenUptime is MIT licensed and there are no check or plan limits when you run it yourself. The shortest path is Docker: one container for the API, the web app and the scheduler, plus a Postgres from the compose file. Prefer Cloudflare? Deploy the Workers to your own account. Either way, the self-hosting guide walks through it.自托管:OpenUptime 以 MIT 协议开源,自己部署没有检查次数和套餐限制。最短路径是 Docker:一个容器跑 API、网页和调度器,再加 compose 文件里的 Postgres。更想用 Cloudflare?把 Worker 部署到你自己的账号就行。两种方式都可以参考自托管文档。
What it does not do (yet)它还做不到的事
We would rather tell you now than have you find out in an incident. There is no SSL certificate expiry monitoring, because the Workers runtime does not expose certificates to our code. Checks run from one place, not from several regions at once. Status pages live on our domain; custom domains are on the list. These are plans, not promises with dates, and we will write about each one when it actually ships.与其让你在故障里才发现,不如现在就说清楚。目前没有 SSL 证书到期监控,因为 Workers 运行环境不向代码暴露证书。检查只从一个位置发出,不是多个地域同时探测。状态页用的是我们的域名,自定义域名在计划里。这些都只是计划,没有承诺日期,真正上线时我们会再写文章介绍。
Go poke it去戳戳看
Add one monitor for something you care about, point an alert at your phone, and let it run for a day. If it ever does the one job it has, and wakes you up before your customers do, it has earned its keep. If something is rough, missing or just plain annoying, tell us in the repository. We read it.给你在乎的东西加一个监控,把告警接到手机上,让它跑一天。哪天它做成了那唯一的任务,赶在客户之前把你叫醒,它就算没白活。如果哪里粗糙、缺失或者单纯让你不爽,来代码仓库告诉我们,我们会看的。