OpenUptime

← All posts全部文章

Why Postgres and not D1 (the short answer is Docker)为什么用 Postgres,而不是 D1(简短回答:Docker)

OpenUptime runs on Cloudflare Workers but keeps its data in Postgres. Four reasons, an honest list of what it costs, and why D1 is not ruled out.OpenUptime 跑在 Cloudflare Workers 上,数据却放在 Postgres。四个理由,一份老实的代价清单,以及为什么没有把 D1 彻底排除。

The question we get most often goes like this: "OpenUptime runs on Cloudflare Workers, and Cloudflare has its own SQL database called D1. Why on earth is there a Postgres in the middle?" Fair. It is the first thing we would ask too.我们被问得最多的一个问题是:"OpenUptime 跑在 Cloudflare Workers 上,Cloudflare 自己就有 SQL 数据库 D1,你们为什么偏要在中间塞一个 Postgres?"问得好,换我们也会先问这个。

The short answer is "Docker". The longer answer is below, including the parts where Postgres makes our life harder. We are not going to pretend it is a free lunch.简短的回答是:"Docker"。长一点的回答在下面,包括 Postgres 让我们日子变难的那几处。我们不打算假装它是顿免费的午餐。

Reason one: it has to run without Cloudflare原因一:它得能在没有 Cloudflare 的地方跑

OpenUptime is something you can host yourself, and "yourself" does not always mean "on Cloudflare". Plenty of people have a small server, a NAS or a Kubernetes cluster, and just want docker compose up. In that setup there is no Worker, no queue and no D1. There is a container and a database.OpenUptime 是可以自己部署的,而"自己"不一定意味着"在 Cloudflare 上"。很多人手边有一台小服务器、一台 NAS 或一套 Kubernetes,只想敲一句 docker compose up。在那种环境里,没有 Worker,没有队列,也没有 D1,只有一个容器和一个数据库。

D1 is SQLite under the hood and lives inside Cloudflare. A database that only exists there cannot be the storage of a product that also has to run in a plain container. Postgres can: the same code talks to the same kind of database in the compose file, on a managed provider, or behind Cloudflare's Hyperdrive. One storage layer, every way of running it.D1 底层是 SQLite,并且只存在于 Cloudflare 内部。一个只能存在于那里的数据库,没法成为同时也要在普通容器里运行的产品的存储。Postgres 可以:同一份代码,既能连 compose 文件里的数据库,也能连托管服务商,还能通过 Cloudflare 的 Hyperdrive 连接。一套存储层,覆盖所有运行方式。

Reason two: the data is the product原因二:数据本身就是产品

A monitoring tool is mostly a machine for writing down what happened, and then being believed later. Check history, incident timelines and uptime numbers are the whole point. So we wanted the boring, battle-tested toolbox around it: psql, pg_dump, point-in-time recovery, read replicas, and a dashboard your team already knows how to read.监控工具本质上是一台"记录发生过什么,并且事后让人信服"的机器。检查历史、故障时间线、可用率数字就是全部价值所在。所以我们想要围绕它的那套无聊但久经考验的工具箱:psql、pg_dump、时间点恢复、只读副本,以及你的团队本来就会看的面板。

It also means leaving is easy. Your data sits in a database you own, in a format every tool understands. If you ever want to take it somewhere else, you do not need our permission or an export button.这也意味着离开很容易。你的数据就在你自己的数据库里,格式所有工具都认识。哪天你想把它搬去别处,不需要我们批准,也不需要什么导出按钮。

Reason three: a minute is a very busy minute原因三:一分钟其实很忙

Every minute the scheduler wakes up, finds the monitors that are due, and hands them out. Picture a deli counter with a ticket machine: several workers may reach for the next ticket at the same moment, and nobody should serve the same customer twice. Postgres has a neat answer for this, FOR UPDATE SKIP LOCKED: each worker grabs the rows nobody else has claimed and skips the ones already taken.调度器每分钟醒一次,找出到点的监控并分发下去。想象一个带取号机的熟食柜台:好几个店员可能同时去拿下一个号,而且同一位顾客不该被服务两次。Postgres 对此有个很漂亮的答案:FOR UPDATE SKIP LOCKED。每个店员只拿还没人认领的行,别人已经拿走的就跳过。

That is how overlapping runs never double-check the same monitor, and also why you can start several containers against one database and have them share the work. Results, incidents and state are written by many checks finishing at about the same time, which is exactly the kind of traffic a database built for concurrent writers handles without drama.这就是为什么重叠的调度周期不会重复检查同一个监控,也是为什么你可以对着同一个数据库启动多个容器,让它们分担工作。检查结果、故障事件和状态,是由很多个差不多同时完成的检查写进去的,这正是为并发写入设计的数据库不费力就能应付的流量。

Reason four: history gets big, graphs must stay cheap原因四:历史会变大,图表得保持便宜

One monitor checked every minute is about 43,000 results a month. Multiply by a few hundred monitors and you do not want every uptime graph to scan raw rows. So statistics come from an hourly rollup table. Raw results are kept for 90 days by default (you can change it) and the rollups for a year. That is a very ordinary job for SQL: group by hour, aggregate, keep it indexed.一个每分钟检查一次的监控,一个月大约产生 4.3 万条结果。乘上几百个监控,你就不会想让每张可用率图都去扫原始记录。所以统计数据来自按小时汇总的表。原始结果默认保留 90 天(可以调整),汇总数据保留一年。这是 SQL 最拿手的活儿:按小时分组、聚合,再配上索引。

Okay, so what does it cost us?好吧,那代价是什么?

Here is the honest list.下面是老实的清单。

  • You need a Postgres. With D1 the setup would be "a Cloudflare account and that is it". With Postgres you also need a database from somewhere. Plenty of hosted ones have a free tier, but it is one more thing to create.你需要一个 Postgres。用 D1 的话,部署就是"一个 Cloudflare 账号,完事"。用 Postgres,你还得从某处弄一个数据库。很多托管服务有免费额度,但终究是多一件要创建的东西。
  • Distance. A Worker may be far from your database, and every query is a round trip. A request makes several of them, so we put the API Worker next to the database on the hosted service, and the self-hosting guide has you create a Hyperdrive to pool and cache connections. D1 sits right next to the Worker and does not have this problem.距离。Worker 可能离你的数据库很远,每条查询都是一次往返,而一个请求会发出好几条。所以在托管版里我们把 API Worker 放在数据库旁边,自托管文档也会让你创建 Hyperdrive 来做连接池和缓存。D1 就在 Worker 旁边,没有这个问题。
  • Connections are not free. Databases only allow so many at once, and Workers can scale up fast. So the code opens one connection per request and closes it afterwards, and the minute cron has to be economical with how it talks to the database.连接不是免费的。数据库同时能接受的连接数有限,而 Workers 可以很快扩容。所以代码里每个请求只开一个连接、用完就关,每分钟的定时任务在访问数据库时也得精打细算。
  • Small-scale limits stay small. On Cloudflare's free Workers plan there is no queue, so the minute cron runs up to 40 monitors inline. That is fine for a handful, and the paid plan with a queue is where bigger setups belong.小规模有小规模的上限。Cloudflare 免费 Workers 套餐没有队列,每分钟的定时任务最多直接跑 40 个监控。几个监控没问题,规模更大就该用带队列的付费套餐。

We left the door open我们把门留了条缝

The application never talks to Postgres directly. It talks to a small storage interface (packages/repo), and Postgres is one implementation of it (packages/db-pg). If a good D1 implementation ever makes sense, it is a new package, not a rewrite.应用从不直接和 Postgres 打交道,它面对的是一层很薄的存储接口(packages/repo),Postgres 只是其中一个实现(packages/db-pg)。如果哪天做一个 D1 实现真的合理,那是新增一个包,而不是重写。

To be clear: a D1 backend does not exist today, and we will not announce one until it does. If a "Cloudflare account and nothing else" setup would make your life easier, tell us in the repository, ideally with what you would run it for. Real use cases are the best way to decide whether to build it.说清楚:现在没有 D1 实现,在它真正存在之前我们不会提前宣布。如果"只要一个 Cloudflare 账号,别的什么都不用"能让你的日子好过一些,请到代码仓库告诉我们,最好说明你打算拿它来做什么。真实的使用场景,是决定要不要做它的最好依据。

New here? The introduction post covers what OpenUptime does, and the self-hosting guide shows both ways to run it.刚来?介绍文章讲了 OpenUptime 能做什么,自托管文档则展示了两种运行方式。