NearBuild runs on one small server: Cloudflare Tunnel, then Caddy, then the web container, and a cron job that syncs the Auckland Council consent PDF every hour. Every health check I had also lived on that server. Last week I decide to fix the obvious problem with that: when the server dies, the thing that should tell me is dead too.
There was a second, quieter gap. When the hourly sync fails after it reaches the database, it records a failed attempt and the public page shows the source as degraded. But some failures never reach the database: cron is not running, the database is unreachable, or the wrapper script exits early. Those write nothing, so the public status stays "healthy" for up to 14 days, until the stale-data rule catches it. The admin page warns after 3 hours, but only if someone opens it.
So I built a small watchdog on AWS Lambda. It does not live on the server, it holds no NearBuild secrets, and it only reads one public URL.
Why Lambda, and why only the watchdog
My first idea was bigger: move the hourly sync itself to Lambda. It does not fit. A normal sync takes 15 to 20 minutes, mostly PDF parsing plus about 5,800 small transactions, and Lambda has a hard limit of 15 minutes. Splitting it would add a queue and more moving parts for no user benefit.
The watchdog is the opposite. It is small job that runs for about one second, needs no state and no database access, and must run somewhere else. That is exactly what Lambda is good at:
| Need | How it is met |
|---|---|
| Independent of the server | Runs in a separate AWS account, in Sydney (ap-southeast-2) |
| Runs on a schedule | EventBridge rule, rate(5 minutes) |
| No secrets | Reads the public /api/public/source-health endpoint only |
| Remembers state between runs | It doesn't. CloudWatch alarms keep the state. |
| Tells a human | Alarm → SNS topic → email, for both ALARM and OK |
The architecture
EventBridge rate(5 minutes)
→ Lambda (Node.js 24, arm64, 256 MB)
GET https://nearbuild.nz/api/public/source-health (10 s timeout, no redirects)
print one EMF log line
→ CloudWatch metrics NearBuild/Watchdog { SiteUp, SourceHealthy, ResponseTime }
→ CloudWatch alarms site-down, source-unhealthy
→ SNS topic → email
The function never decides when to send an alert. It only measures and prints. The alarms decide. This keep the function stateless, and it means a crashed function is also visible: no data is a signal too.
What one check does
The check makes one bounded GET and turns every failure into data instead of an exception:
const response = await fetch(url, {
headers: { accept: "application/json", "user-agent": "NearBuildWatchdog/1.0 (+https://nearbuild.nz)" },
redirect: "manual", // a redirect is a failure, not something to follow
signal: AbortSignal.timeout(timeoutMs), // also covers reading the body
});
Then a pure function evaluates the result. It has two outputs:
| Metric | Value is 1 when |
|---|---|
SiteUp | HTTP 200 and a well-formed source-health payload |
SourceHealthy | the site is serving the live database (not the bundled snapshot), the status is healthy, and the last sync attempt is at most 3 hours old |
The 3-hour rule is the important one. It closes the gap from the start of this post: if cron stops, the status may still say healthy, but lastAttemptAt stops moving, and the watchdog sees it. The number is the same as the admin warning in the web app, and a comment in both places says to keep them in step.
The evaluator is plain TypeScript with no AWS code in it, so it is easy to test. There are 24 tests for timeouts, non-JSON bodies, redirects, the snapshot fallback, the exact 3-hour boundary and a source that never synced.
Metrics without an API call
The usual way to publish a custom metric is cloudwatch:PutMetricData. I used CloudWatch Embedded Metric Format (EMF) instead. The function prints one JSON line, and CloudWatch extracts the metrics from the log stream:
{"_aws":{"Timestamp":1790799468495,"CloudWatchMetrics":[{"Namespace":"NearBuild/Watchdog",
"Dimensions":[["Service"]],"Metrics":[{"Name":"SiteUp","Unit":"Count"},
{"Name":"SourceHealthy","Unit":"Count"},{"Name":"ResponseTime","Unit":"Milliseconds"}]}]},
"Service":"nearbuild","SiteUp":1,"SourceHealthy":1,"ResponseTime":2528,
"hoursSinceAttempt":0.55,"reasons":[]}
It don't need any extra IAM permission, and the reasons field stays searchable in the logs without becoming a metric. It is also much more simpler than calling an API on every run. The function's role can only write to its own log group. That is the whole policy.
The alarms
An alarm that fires on one bad check is noise. An alarm that waits too long is useless. I picked the windows from how NearBuild really behaves:
SiteDownAlarm:
MetricName: SiteUp
Statistic: Maximum # 1 if any check in the period passed
Period: 600 # about two checks per period
EvaluationPeriods: 3 # 30 minutes
Threshold: 1
ComparisonOperator: LessThanThreshold
TreatMissingData: breaching # the watchdog stopping is also an alert
- site-down: every check failed for 30 minutes.
Maximummeans one good check in a 10-minute period is enough to call it up, so a single slow request does not page me. - source-unhealthy: 3 periods of 30 minutes, so 90 minutes. One failed hourly sync makes the source degraded until the next run succeeds, which can be about an hour. I don't want an email for one failure that fixes itself. Two failures in a row, a stuck run or a stopped scheduler is different.
- Both alarms treat missing data as breaching. Every alarms also sends OK, so I know when it recovers.
Deploying it with the Serverless Framework
The whole stack is one serverless.yml: the function and its schedule at the top, and the SNS topic, email subscription and two alarms as raw CloudFormation under resources. The framework turns it into a CloudFormation stack with 13 resources, and the upload is one file of 4.6 KB, because the handler is bundle with esbuild and everything else is excluded.
Choosing the version take longer than writing the handler:
- Serverless Framework v3 refused my config. Its schema only knows runtimes up to
nodejs20.x, and new functions on Node 20 are no longer a good idea:Configuration error at 'provider.runtime': must be equal to one of the allowed values. - Serverless Framework v4 supports new runtimes but requires a login to their dashboard, even for a free personal project.
- osls, the open-source fork of v3, supports
nodejs22.xandnodejs24.xwith the same config format and no login. I use this one. Under the hood it is still CloudFormation, which I already know well.
The watchdog is also not an npm workspace in the NearBuild repo. It has its own lockfile, and the folder is in .dockerignore, so its tooling can never end up in the web or worker images.
Least privilege for the deploy, too
The function's role was easy. The deploy identity needs more care, because a deploy creates roles, buckets and functions. I wrote a policy that only allows the actions this stack needs, and only on resources named nearbuild-watchdog-prod-*: the CloudFormation stack, its deployment bucket, the function, the function's role, its log group, the schedule rule, the SNS topic and the two alarms. It is attached to a role, not a user, so there is no long-lived access key on my machine. I log in with aws login, which gives temporary credentials, and the deploy assumes the role.
IAM Access Analyzer found 0 issues in the policy. Then I checked both directions with the real role:
| Action through the deploy role | Result |
|---|---|
| Deploy a real change to the stack | Allowed, 23 seconds |
| Invoke the watchdog function | Allowed |
| List S3 buckets | Denied |
| List IAM users / create an IAM user | Denied |
| List other Lambda functions | Denied |
| Describe EC2 instances | Denied |
What the first deploy taught me
Five things surprised me on the first day, and one of them I caused on purpose.
1. A new alarm starts in ALARM
Two minutes after the first deploy, the site-down alarm went from INSUFFICIENT_DATA to ALARM. Nothing was broken. The alarm was brand new, there was no data yet, and I had told it that missing data means failure. It went back to OK two minutes later, when the first metric arrived. The email subscription was not confirmed yet, so nobody got a false alert, but next time I deploy a new alarm I expect this and confirm the subscription after the first data point.
2. 128 MB was too small
The first run used 103 MB of memory, and a later one used 123 MB, out of 128 MB. A Node.js 24 runtime plus fetch does not leave much room. I raised it to 256 MB. It is still inside the free tier, and more memory on Lambda also means more CPU.
3. Least privilege finds its own gaps
During the drill below I wanted to read the alarm history through the deploy role, and it was denied: I had given it DescribeAlarms but not DescribeAlarmHistory. That is the correct failure direction. I added the one read-only action, validated the policy again, and published a new policy version.
4. Breaking it on purpose
An alarm that never fired is only a theory. So I deployed the watchdog with its URL pointing at a path that returns 404, and watched:
| Time (UTC) | What happened |
|---|---|
| 20:21:28 | Deployed the broken URL. Every check now reports SiteUp = 0. |
| 20:51:46 | site-down went to ALARM, 30 minutes and 18 seconds later, and SNS published the alert |
| 21:02:49 | Deployed the correct URL again |
| 21:07:46 | site-down went back to OK, and SNS published the recovery |
Thirty minutes is what I designed: three 10-minute periods where every check failed. Five minutes after the fix, one good check was enough to turn it back to OK.
5. The alert channel can be switched off by one click
This was the most useful lesson of the day. Soon after I confirmed the email subscription, AWS sent an "Unsubscribe Confirmation": the subscription was gone. SNS email confirmations are not authenticated by default, so every alert email and the confirmation page carry an unsubscribe link that works for anybody who opens it. A stray click, a mail client's one-click unsubscribe or a link scanner can all remove it, and CloudTrail does not log these link-based unsubscribes, so I cannot even see which one it was.
Then it got stranger. I recreated the subscription with CloudFormation, and SNS gave back the same subscription ID, still deleted, and sent no new confirmation. SNS reuses a recently deleted subscription for the same address. CloudFormation reported success while the only alert channel was dead. I moved the alerts to a new address (a Gmail plus-address, which also makes them easy to filter), and the rule now is: confirm through the API with AuthenticateOnUnsubscribe, never by clicking the link, so an unsubscribe must be a signed AWS request. I also check the SNS delivery metrics after any change to the channel, because "CREATE_COMPLETE" said nothing about whether an email can arrive.
Cost
| Item | Per month | Free tier |
|---|---|---|
| Lambda invocations (every 5 minutes) | 8,640 | 1,000,000 requests |
| Lambda compute (about 1.5 s at 256 MB) | about 3,300 GB-s | 400,000 GB-s |
| Custom metrics / alarms | 3 / 2 | 10 / 10 |
| SNS email notifications | a few | 1,000 |
So it should cost nothing. I still added a $5 monthly budget for warn me by email at 80%, because "should cost nothing" is exactly the sentence before a surprise bill.
Limits I accept
- It checks from one place, Sydney. Although a break only between Sydney and Cloudflare would give me a false alert, but for a beta product that is fine.
- It says "down", not why. The tunnel, Caddy and the web container all look the same from outside. The alarm text lists where to look, in order.
- It reads the same endpoint the users read. If that endpoint lies, the watchdog believes it. That is why the endpoint's own rules are tested separately in the web app.
The lesson is simple. A health check that shares the machine with the thing it checks can only tell you about the problems the machine survives. Put the watcher somewhere else, keep it small, and let the alarms, not the code, decide when to wake you.