My team runs a group of small content sites. Each one is a pre-built front end: HTML, CSS, JavaScript and some JSON files. Articles come from a shared content API, and each site's Nginx proxied /api/ to that API. For years they lived on one busy shared server that carries most of the company's revenue.
That server was a single point of failure:
- It runs about 630 containers, and the rule is it must never go down.
- The shared Nginx had a 256 MB memory cap. One day the master process was killed during a reload, and every site on the machine was offline for 7 seconds.
- A previous static host had already died once, with no copy anywhere else.
These sites is pre-built files. They don't need a server at all. So in September we move them to AWS: CloudFront in front of private S3 buckets, one CloudFormation stack per site, with Cloudflare still in front for DNS, WAF and caching. This post is the full route, with the real numbers and the mistakes.
First, what we did not do
The first idea was to move the whole server to AWS. A cost study killed it quickly:
| Option | Monthly cost |
|---|---|
| The whole current server (rented dedicated box) | about $520 |
| Only the egress on AWS, about 6 TB a month | about $531 |
Metered bandwidth alone cost more than the whole box. We also tried a plain EC2 lift-and-shift for the static sites: 90 of 93 went live, and the next evening 97 DNS records were pointed back to the old server. Moving files to another server did not remove the server problem. Static files belong on object storage and a CDN.
Target architecture
visitor → Cloudflare (orange cloud, WAF and rules kept, edge cache)
→ CloudFront (HTTPS, second cache layer)
→ Origin Access Control (SigV4, always)
→ private S3 bucket: releases/<release-id>/
browser → content API (unchanged, verified separately)
The old server stays alive as the rollback origin until each site is proven. Fixed choices:
- S3: the REST regional endpoint, not the website endpoint, because the website endpoint does not work with OAC. All four Block Public Access settings on, versioning on, and the bucket policy allows exactly one distribution.
- Immutable releases: files go into
releases/<release-id>/, and a new version is a new prefix, never an overwrite. Uploads usePutObjectwithIf-None-Match: *. - Verification by SHA-256, never the S3 ETag, because multipart uploads and encryption change the ETag.
- Certificate: ACM in us-east-1 (CloudFront only accepts that region), DNS-validated, exact hostname only.
- Shared policies: the default quotas allow 500 distributions but only 20 custom cache policies, so a policy per site will not fit. One shared cache policy (min 0, default 60, max 300 seconds, gzip and Brotli, no cookies or headers in the key, only the
vquery parameter) and two header policies.
One CloudFormation stack per site
Each site has own stack with exactly four resources:
Resources:
SiteBucket # Retain on delete and replace
OriginAccess # OAC: s3 / always / sigv4
Distribution # DefaultRootObject index.html, origin path /releases/${ReleaseId}
BucketPolicy # s3:GetObject for this distribution ARN only, deny non-TLS
Why a stack per site:
- Isolation. A mistake on one site cannot touch another.
- Reviewable change sets. Before anything runs, the change set must show only these four resources and no replacement.
- Per-site rollback of DNS, release and template.
- Tag-scoped IAM. Stack names follow a pattern with a project tag, so the permission policy can say "only these stacks".
Stack names are a slug of the hostname plus the first 8 hex characters of a SHA-256 of the full hostname. The hash is computed before the slug is cut to 35 characters, so two long hostnames can never produce the same name.
Sites with an API need a small router
CloudFront's default root object only applies to /, so /about/ will not become /about/index.html by itself. And many sites used the same-origin /api/. For them the stack adds a second origin (the content API, HTTPS only), two behaviours for /api and /api/* with caching disabled, and a CloudFront Function of about 7.7 KB:
function handler(event) {
var request = event.request;
if (request.uri === '/api') return redirect('/api/'); // keep old trailing-slash behaviour
if (request.uri.indexOf('/api/') === 0) {
request.uri = '/wp-json/wp/v2/' + request.uri.slice(5); // rewrite to the content API
return request;
}
var path; try { path = decodeURIComponent(request.uri); } catch (e) { path = ''; }
if (known.hasOwnProperty(path)) return request; // file exists in the manifest
return notFound(); // original 404.html bytes, no-store
}
We learned this after the fact. At first the distribution mapped every 403 and 404 to /404.html. That mapping also turned the API's JSON errors into HTML, and on 29 migrated sites /api fell into the static 404. Home pages looked fine, but articles didn't load. The router fixed all 29, and each repair round re-verified all 3,766 static files.
The per-site flow
The pilot was the site with the fewest requests: 10,135 in six days, 130 public files, 8,490,476 bytes. Its DNS cutover happened on 17 September at 00:25:58 UTC. After that, sites went in batches of up to five, and a later serial run of 44 sites took about 13 to 20 minutes per site. Every site goes through the same steps:
- Freeze and baseline. Save the DNS record (ID, proxy status, TTL, comment, tags), Nginx rules, the 404 page and current cache headers. Screenshot the home page, two real articles, a category page and the mobile menu. Build a manifest of every public file with path, size, SHA-256 and MIME type, rejecting hidden files, symlinks and anything like
.sql,.pem,.bakor.zip. This manifest keeps all the informations we need for verification and rollback. - Candidate stack. Validate the template, create a change set limited to the four resource types, review it, execute it, and wait for
CREATE_COMPLETEand aDeployeddistribution. An accepted API call is not the same as ready. - Upload and verify. Conditional uploads into the new release prefix. Fetch every file through the CloudFront candidate URL and compare body hash and MIME type. Anonymous S3 access must return 403. A repeat request must show
Hit from cloudfront. Download one version set into a separate folder and hash it: that is the restore drill. - Browser check. The page must show real API content, like an article body and a category title. A page shell that returns 200 does not count.
- Certificate. Request in us-east-1 with DNS validation, add the validation CNAME as grey cloud and keep it for renewals, then an update change set adds the alias and certificate.
- Formal pre-check with the real hostname and SNI going straight to CloudFront.
curl -kis banned.curl --connect-to '<host>:443:<dist>.cloudfront.net:443' https://<host>/ - DNS cutover. PATCH the same Cloudflare record ID, changing only the type and content (A record to CNAME), keeping proxy status, TTL, comment and tags. Read it back and compare. Purge the exact URLs from the manifest, never the whole zone.
- Public verification. Every file, the home page, a real 404, CloudFront headers like
ViaandX-Amz-Cf-Id, the browser checks again, S3 still 403, and neighbouring sites still on the old origin.
Rollback is automatic if key pages fail twice, content is wrong, permissions leak, an unexpected noindex appears, or the cache probe fails. It only runs if the live record still matches what we wrote, and it restores exactly the record we saved.
Incidents along the way
The release switch that returned 403
To publish new header images on 81 sites, we changed the bucket policy and the origin path in one CloudFormation update. The bucket policy references the distribution, so CloudFormation update the distribution first, which takes about three minutes, and the policy after. In that gap the edges read the new prefix and got AccessDenied, and even the custom 404 page could not be read.
| First 49 sites | 403 responses |
|---|---|
| Same hour, the day before | 14 |
| During the switch | 292 (about 230 from real visitors) |
The fix is a two-phase switch, and the two steps are never in one change set:
- Widen the bucket policy to allow both the old and the new release.
- Switch the origin path and wait for
Deployed. - Narrow the policy to the new release.
Smaller ones
- Flexible SSL. One host had a Cloudflare rule with SSL mode Flexible, so Cloudflare talked to CloudFront over HTTP, CloudFront redirected to HTTPS, and the cutover rolled back automatically. The fix was Full SSL for that host, not "make the error go away" with Flexible.
- Email obfuscation. Cloudflare rewrites email addresses in HTML, so 3 of 130 pages per site did not match the raw hash. We decode the email and remove the injected script before hashing, instead of turning the protection off.
- Zone guessing. Taking the last two labels of a hostname broke on a
.co.nzdomain. The zone now comes from the API. - A correct first MISS. One site rolled back because its cache probe got a normal first MISS. The probe now allows up to 3 warm-up samples but still requires a real HIT.
- Line endings. A generator changed the 404 page from CRLF to LF. The byte-exact check caught it and rolled back.
The bill: requests, not bandwidth
People usually think bandwidth is the expensive part of a CDN. For us it was not. In the same time we were migrating, I checked Cost Explorer. AWS for September up to the 28th was about $619 net, and CloudFront was $130.57 of it:
| CloudFront item | Volume | Cost |
|---|---|---|
| US HTTPS requests | 125.3M | $118.31 |
| CloudFront Functions executions | 113.3M | $11.13 |
| Canada HTTPS requests | 1.11M | $1.10 |
| Data transfer out | about 215 GB | $0 (free tier) |
The average response was only about 1.5 KB, so the requests is what we paid for. And the daily cost was climbing fast:
| Day | CloudFront cost |
|---|---|
| 24 September (free tier of 10M requests ran out) | $11.18 |
| 25 September | $33.93 |
| 26 September | $34.50 |
| 27 September | $39.27 |
That pace is about $1,000 to $1,200 a month. CloudWatch showed requests across all 276 distributions went from 1.58M a day on 23 September to 36.89M a day on 27 September, more than twenty times, after two day of migration batches. That is much more than the traffic growth.
So I probed every host and read the cf-cache-status header:
| Hosts | Count | Share of requests (27 Sep) |
|---|---|---|
HTML returned DYNAMIC (not cached by Cloudflare) | 131 | 93.9% (34.63M) |
| Cached normally | 144 | 6.1% |
The 131 hosts were about 90 migrated static sites plus sites that our CMS publishes straight to S3, and all of them were on the same Cloudflare rule. The busiest single site made 1.3 to 2.85 million requests a day.
It was our own design. At the start I decided CloudFront should be the only cache layer, so invalidation lives in one place, and every migrated host was added to a Cloudflare rule with caching turned off. It sound clean. In practice every visitor request went through Cloudflare to CloudFront and ran the router function, and CloudFront billed each one. The function alone was worth about $100 a month.
The fix
21 Cloudflare zones, 29 operations (18 rule updates, 8 new rules, 3 new cache-phase entrypoints):
- The old "no cache" rules were narrowed to only
/apiand/api/*. The API must stay uncached, because the upstream sendscache-control: max-age=16070400, about six months. - 11 zones without a zone-wide "cache everything" rule got a host-scoped rule that caches everything except
/api. - Rule order matters: in the cache phase, a later matching rule overrides an earlier one.
"s3go hosts: Cloudflare cache (except /api)" {"cache": true}
(http.host in {...}) and not (http.request.uri.path eq "/api"
or starts_with(http.request.uri.path, "/api/"))
"s3go hosts: bypass /api only" {"cache": false}
(http.host in {...}) and (http.request.uri.path eq "/api"
or starts_with(http.request.uri.path, "/api/"))
The script took the same locks as the migration, saved the full rulesets as a rollback baseline, did a dry run to prove no host outside the list was touched, and stopped on the first error.
Verification and effect
| Check | Result |
|---|---|
| Home pages, first request | 108 MISS, 17 HIT, 6 REVALIDATED |
| Home pages, second request | 131/131 HIT, all HTTP 200 with CloudFront headers |
/api | 131/131 DYNAMIC (81 returned 200, 50 never had an API route) |
| All distributions, requests per 5 minutes | 135k to 170k before, 35k to 48k after: about 27% |
| The 131 hosts alone | 16% of their previous requests |
| One example site, per 20 minutes | 53,000 to 4,450 requests |
At that rate, CloudFront should go from about $39 a day to about $11 a day, roughly $800 a month saved. This is still a projection from the first hour; I will confirm it with a full day of billing data. The next lever is raising HTML s-maxage from 60 seconds to 5 or 10 minutes, since every release already purges exact URLs.
The process changed too. HTML returning DYNAMIC is now an acceptance failure, and new whole-host "no cache" rules are banned. One side note for anyone doing the same: the zones' Browser Cache TTL of 14,400 seconds rewrites the visitor header to max-age=14400, so browsers may keep old HTML for up to four hours, and a CDN purge cannot clear that.
Least privilege for a second operator
Later another person need to publish sites without full AWS access. They log in with IAM Identity Center and MFA, and get an inline-only permission set with a four-hour session:
| Allowed | Condition |
|---|---|
| Create change set | Stack name pattern, project tag, only the 4 resource types, no execution role, no resource import |
| Upload objects | Only releases/*, only with if-none-match (so plain aws s3 sync fails, on purpose) |
| Create / update distribution | Only through CloudFormation, only with the project tag |
| Request certificate | Only us-east-1, only DNS validation, only with the project tag |
| Not allowed | Deleting buckets, objects, versions or distributions; changing shared cache policies; any IAM action |
Validation: AWS Access Analyzer found 0 issues, 26 of 26 simulated allow and deny cases passed, and 5 more simulations against the real generated role denied iam:PassRole and ec2:TerminateInstances. An audit of the 37 existing stacks found none using an execution role, which supports the "no role" condition.
What I would repeat
- Treat every site as its own small, reviewable unit.
- Verify bytes, not status codes.
- Make rollback a tested path that restores exactly what you saved.
- After a migration, read the bill by request count, not only by bandwidth.
Our biggest cost problem was not AWS pricing. It was one caching decision that looked clean on paper.