PayTaxNZ is a tax tool for self-employed people and freelancers in New Zealand. You record invoices, income, expenses and assets, and it calculates GST and income tax for you. The first version was a Laravel 10 app with Dcat Admin and MySQL. It worked, and people used it for real tax filing. In August I start to rewrite it as a TypeScript monorepo, and the first production deploy of the new system was nine days later.
Rewriting a tax system is scary. If a button looks wrong, users tell you. If a depreciation number is wrong by 50 dollars, nobody tells you until the tax office does. So the whole project was designed around one question:
How do I prove the new numbers are right?
Why rebuild instead of refactor
People always ask me why don't I refactor the old code. The honest answer is that the tax engine inside the old system was good, but three structural problems could not be fixed step by step:
- Tenant isolation leaked. It depended on an implicit global function,
admin_user(), called in 144 places. In a queue job or a CLI command this global is empty. Two internal audits found leaks, and each fix was only a patch. - Tests did not use the real schema. Test tables were written by hand inside the tests, so a passing test told you very little about production.
- No growth path. No payments, no self sign-up, no API and no mobile app. Almost half of the 18,000 lines of PHP were admin CRUD boilerplate.
So I kept the ideas and replaced the frame. The old system have some very good ideas, and all of them moved to the new design:
- a four-layer chain for assets: facts, rules, a yearly snapshot, then a summary
- IRD rules stored as data, not hard-coded
- per-company setting overrides
- audit fields on every business table
- every bug ever fixed in the old code becomes a regression test in the new code
A modular monolith, not microservices
We discussed about microservices for maybe ten minutes. A tax summary is basically one big join: income, expenses, categories, depreciation snapshots, pool snapshots and disposal gains. If you split the database, you lose transactions. In tax, "eventual consistency" means there is a window where the user sees a wrong filing number. That is not acceptable.
So it is one database and one codebase, with three extra processes where isolation really helps:
| Process | Why it is separate |
|---|---|
worker (pg-boss) | Tax recalculation must not compete with web requests |
pdf (Playwright/Chromium) | Chromium memory use and crashes stay away from the web app |
ai | A slow, failure-prone external API gets its own rate limits and retries |
Separate processes are not microservices. The monorepo (pnpm + Turborepo) looks like this, and dependencies only point one way:
packages/domain pure tax engine, only depends on decimal.js
packages/db Drizzle schema, RLS policies, SQL migrations
packages/api oRPC procedures, the only entry point for business logic
packages/auth Better Auth
packages/i18n 8 locales, zero dependencies
apps/web Next.js 15 App Router, never imports db directly
apps/worker pg-boss jobs
apps/pdf Playwright/Chromium
apps/ai receipt extraction and classification
domain ← api ← web / worker / pdf / ai
In production this becomes seven long-running containers (Postgres, MinIO, web, worker, pdf, ai and a backup loop) plus a one-shot migrate service. Every app container depends on migrate finishing successfully, so a failed migration cannot leave a half-started system.
The tax engine is pure, and a test enforces it
packages/domain is pure TypeScript package. A small test scans the source and fails if any production file:
- imports anything except relative files or
decimal.js - uses
Dateornew Date - uses
Math.random
Banning Date sounds extreme, but it solved a real problem. A Date carries a timezone and an instant. Once it enters tax maths, "the same date falls into a different period in a container with a different timezone" becomes possible, and you cannot reproduce it reliably in a unit test. So every tax dates is a plain YYYY-MM-DD string. Pure functions can be tested exhaustively, and later the same code can run on mobile for offline estimates.
Golden assertions, written before the code
Before writing any engine code, I wrote a numbered spec called the golden assertions. Each assertion has an ID like LV-04 or DISP-01, and the same ID is the prefix of the test name, so you can go from spec to test and back.
| Class | Source | Expectation | Count |
|---|---|---|---|
| A. Rules | The old system's specs | New and old agree | 161 |
| B. Regressions | 10 bugs fixed in the old system | New must not repeat them | 17 |
| C. Open defects | An audit of the old system | New must fix them, so new and old disagree on purpose | 8 |
| Total | 186 | ||
This split was important. If I treated class C as the baseline, I would copy old bugs into the new system. If I ignored them, the reconciliation later would show differences that nobody could explain.
The git history shows the order: the spec was the first commit of the repo at 10:36, and the first engine code, the GST period resolver, came at 11:17.
Two habits made the assertions much stronger:
- Pairs. One assertion says there is zero depreciation in the year you sell an asset; its pair says this must not wipe depreciation already accumulated in earlier years. A single assertion can pass while hiding a bug.
- Boundaries. The low-value asset threshold changed several times, so the tests check each boundary day. Testing only middle values is not testing.
// LV-04: low-value asset threshold by purchase date
it.each([
['2020-03-16', '500'],
['2020-03-17', '5000'],
['2021-03-16', '5000'],
['2021-03-17', '1000'],
])('threshold on %s is %s', (date, expected) => { /* ... */ });
I also did mutation checks: break the code on purpose and confirm a test turns red. In one domain, 2 of 7 mutations survived, which showed two blind spots in the tests. I very like this exercise, because it tests the tests.
The test suite grew from 23 tests at the start to around 1,000 at the i18n launch, and today it is over 1,100.
Reconciliation: recompute, don't copy
Migration is where most rewrites cheat a little. I tried not to:
- Raw tables are copied (companies, invoices, expenses, assets). Relationships are rebuilt through
id_map(table, legacy_id, new_id). - Derived tables are not copied at all (GST returns, income tax results, depreciation snapshots). The new engine recomputes them from the copied raw data.
- A script compares every company, every tax year and every GST period, field by field, at cent precision.
Copying only proves that two databases agree. Recompute-then-compare is both the migration and the strongest acceptance test.
Each value gets one label: MATCH, KNOWN(<ledger id>) or UNEXPLAINED. The classifier is code, not my opinion, and a known class must reproduce the exact amount of the difference. For example, for the "unpaid invoice residue" class, the income difference must equal the sum of the stripped rows within one cent.
| Result on 1,070 values | Count |
|---|---|
MATCH | 999 |
KNOWN | 71 |
UNEXPLAINED | 0 |
Although the biggest group of known differences looks scary, but it is harmless: 39 values are zero-amount GST periods that the old system created by a bug, and the new system just regenerates them correctly. The full breakdown:
| Class | What it is | Values |
|---|---|---|
| D-03 | Empty GST periods from an old period-generation bug, regenerated | 39 |
| OPEN-05 | Income left behind when a paid invoice was set back to unpaid | 10 |
| D-11 | Old diminishing-value depreciation used a column nobody updated | 9 |
| cascade | Summary fields following an already classified component | 7 |
| REG-03 | Old summary included assets still pending review | 3 |
| STALE | Old summary row never refreshed after changes | 2 |
| D-10 | GST for an unregistered period | 1 |
D-11 is the most serious one. The old code calculated diminishing value as baseCost - accumulated_depreciation, but that column was never updated anywhere and was always 0 in production. So from the second year, the old system charged the rate on the original cost instead of on the reducing balance. On the golden sample the old value is 496.98 and the new value is 443.06.
Reconciliation also found bugs in my new code, which was the whole point:
- Invoice unit prices in the old system exclude GST. My first version treated them as GST-inclusive, even with a comment that said "same as legacy".
- The expense deduction rate was never applied, so every expense was treated as 100% deductible.
- The deduction amount must be (total − GST) × rate, and the first version used total × rate.
None of these were found by reading code. They were found by comparing with real, anonymised data. The rule I follow now: when a difference appears, check the divergence ledger first, then suspect the new engine. And any commit that changes a rule compared to the old system must add a ledger row in the same commit.
Data safety: fail closed
Row-level security
CREATE FUNCTION app.current_company_id() RETURNS uuid LANGUAGE sql STABLE AS $$
SELECT NULLIF(current_setting('app.company_id', true), '')::uuid
$$;
-- created in a loop over every tenant table
USING (company_id = app.current_company_id())
WITH CHECK (company_id = app.current_company_id())
- Each request runs in a transaction that sets
app.company_idwithset_config(..., true), the same asSET LOCAL. - Missing context gives zero rows, not an error and not everything, thanks to
NULLIF. - The app role cannot bypass RLS, and tables use
FORCE ROW LEVEL SECURITY, so even the owner is filtered. - Policies are created in a loop, because copy-pasted policies drift.
- RLS is not enough alone: the company ID comes from a cookie, so every request also checks membership. "Missing" and "not a member" return the same error, so IDs cannot be enumerated.
- Jobs that cross tenants must be declared as system jobs with a written justification.
Money
Money is numeric(18,4) in the database and decimal.js in code. It travels as a string, never a JSON number, and rounding is half-up to match the old PHP round(). An audit still found a bug: an edit form refilled amounts with Number(total).toFixed(2), so 115.1234 silently became 115.12 when you saved without changing anything.
Idempotency and migrations
- Every write needs an idempotency key. It is required, not optional, because the weak-network mobile retry is exactly the path most likely to forget an optional key.
- The key row and the business write are in the same transaction. A SHA-256 fingerprint of the procedure plus the canonical input detects the same key reused with a different body.
- Migrations are expand-and-contract: add, confirm, then delete in a separate step. CI rejects any pull request that edits an existing migration file.
- Because migrations only expand, old code still runs on the new schema, and that is what make rollback possible.
Email and AI, also designed not to lie
Outbox email. Sending an invoice writes a row into an outbox table inside the same transaction as the status change. A worker job runs every minute, claims rows with FOR UPDATE SKIP LOCKED, and sends outside the transaction. So "status changed but email not queued" and "email sent but transaction rolled back" cannot happen. A send that times out is marked failed and not resent, because a customer receiving the same invoice twice is worse.
AI document inbox. One rule: AI only suggests, it never books.
- The model reads a receipt.
- Zod validates the response: amounts must be decimal strings, currency must be NZD, an unknown category becomes null.
- Deterministic rules in the domain package decide: expense, asset or personal purchase.
- The user confirms every posting.
One bug from this area was very surprising. The claim query used WHERE id IN (SELECT ... LIMIT n FOR UPDATE SKIP LOCKED). With RLS, the planner chose a nested loop and ran the subquery again for each outer row, so a batch size of 2 claimed 3 rows. It only reproduced on the shared dev database.
-- simplified shape of the fix
WITH picked AS MATERIALIZED (
SELECT id FROM outbound_emails
WHERE status = 'queued'
ORDER BY created_at
LIMIT $1
FOR UPDATE SKIP LOCKED
)
UPDATE outbound_emails SET status = 'sending'
FROM picked WHERE outbound_emails.id = picked.id;
Releases with gates
There were 13 production releases in the first two weeks. Each one follows the same steps:
- Preflight on the host: disk capacity must pass.
- Image tags set to the commit SHA, never a moving tag.
docker compose up -d- Smoke script: web health with retries, the PDF service, the redirect from the old host, storage returning 403 to anonymous users (which proves the request really reached storage), and the backup service running.
- The env file is saved with the previous SHA before every release, so rollback is two image tags and the same smoke test.
Some lessons came from real releases:
- One release pulled new images, but the old compose file was still on the server, so a new environment variable was missing. Now the release compares checksums of local and remote files first.
- A
sedreplacement silently did nothing because the key did not exist yet, and the AI service started disabled. Now environment edits are upserts.
What I would tell myself before starting
- Write the yardstick before the thing you measure.
- Recompute instead of copying, and give every difference a name.
- Make the safe behaviour the default: zero rows when context is missing, required idempotency keys, expand-only migrations.
None of this is fancy. It spend more time at the start, but it is the reason I can say the new numbers are right, and show the evidence.