PayTaxNZ is a tax tool for self-employed people and freelancers in New Zealand. You record invoices, income, expenses and assets, and it calculates GST and income tax for you. The first version was a Laravel 10 app with Dcat Admin and MySQL. It worked, and people used it for real tax filing. In August I start to rewrite it as a TypeScript monorepo, and the first production deploy of the new system was nine days later.

Rewriting a tax system is scary. If a button looks wrong, users tell you. If a depreciation number is wrong by 50 dollars, nobody tells you until the tax office does. So the whole project was designed around one question:

How do I prove the new numbers are right?

Why rebuild instead of refactor

People always ask me why don't I refactor the old code. The honest answer is that the tax engine inside the old system was good, but three structural problems could not be fixed step by step:

  1. Tenant isolation leaked. It depended on an implicit global function, admin_user(), called in 144 places. In a queue job or a CLI command this global is empty. Two internal audits found leaks, and each fix was only a patch.
  2. Tests did not use the real schema. Test tables were written by hand inside the tests, so a passing test told you very little about production.
  3. No growth path. No payments, no self sign-up, no API and no mobile app. Almost half of the 18,000 lines of PHP were admin CRUD boilerplate.

So I kept the ideas and replaced the frame. The old system have some very good ideas, and all of them moved to the new design:

  • a four-layer chain for assets: facts, rules, a yearly snapshot, then a summary
  • IRD rules stored as data, not hard-coded
  • per-company setting overrides
  • audit fields on every business table
  • every bug ever fixed in the old code becomes a regression test in the new code

A modular monolith, not microservices

We discussed about microservices for maybe ten minutes. A tax summary is basically one big join: income, expenses, categories, depreciation snapshots, pool snapshots and disposal gains. If you split the database, you lose transactions. In tax, "eventual consistency" means there is a window where the user sees a wrong filing number. That is not acceptable.

So it is one database and one codebase, with three extra processes where isolation really helps:

ProcessWhy it is separate
worker (pg-boss)Tax recalculation must not compete with web requests
pdf (Playwright/Chromium)Chromium memory use and crashes stay away from the web app
aiA slow, failure-prone external API gets its own rate limits and retries

Separate processes are not microservices. The monorepo (pnpm + Turborepo) looks like this, and dependencies only point one way:

packages/domain   pure tax engine, only depends on decimal.js
packages/db       Drizzle schema, RLS policies, SQL migrations
packages/api      oRPC procedures, the only entry point for business logic
packages/auth     Better Auth
packages/i18n     8 locales, zero dependencies
apps/web          Next.js 15 App Router, never imports db directly
apps/worker       pg-boss jobs
apps/pdf          Playwright/Chromium
apps/ai           receipt extraction and classification

domain ← api ← web / worker / pdf / ai

In production this becomes seven long-running containers (Postgres, MinIO, web, worker, pdf, ai and a backup loop) plus a one-shot migrate service. Every app container depends on migrate finishing successfully, so a failed migration cannot leave a half-started system.

The tax engine is pure, and a test enforces it

packages/domain is pure TypeScript package. A small test scans the source and fails if any production file:

  • imports anything except relative files or decimal.js
  • uses Date or new Date
  • uses Math.random

Banning Date sounds extreme, but it solved a real problem. A Date carries a timezone and an instant. Once it enters tax maths, "the same date falls into a different period in a container with a different timezone" becomes possible, and you cannot reproduce it reliably in a unit test. So every tax dates is a plain YYYY-MM-DD string. Pure functions can be tested exhaustively, and later the same code can run on mobile for offline estimates.

Golden assertions, written before the code

Before writing any engine code, I wrote a numbered spec called the golden assertions. Each assertion has an ID like LV-04 or DISP-01, and the same ID is the prefix of the test name, so you can go from spec to test and back.

ClassSourceExpectationCount
A. RulesThe old system's specsNew and old agree161
B. Regressions10 bugs fixed in the old systemNew must not repeat them17
C. Open defectsAn audit of the old systemNew must fix them, so new and old disagree on purpose8
Total186

This split was important. If I treated class C as the baseline, I would copy old bugs into the new system. If I ignored them, the reconciliation later would show differences that nobody could explain.

The git history shows the order: the spec was the first commit of the repo at 10:36, and the first engine code, the GST period resolver, came at 11:17.

Two habits made the assertions much stronger:

  • Pairs. One assertion says there is zero depreciation in the year you sell an asset; its pair says this must not wipe depreciation already accumulated in earlier years. A single assertion can pass while hiding a bug.
  • Boundaries. The low-value asset threshold changed several times, so the tests check each boundary day. Testing only middle values is not testing.
// LV-04: low-value asset threshold by purchase date
it.each([
  ['2020-03-16', '500'],
  ['2020-03-17', '5000'],
  ['2021-03-16', '5000'],
  ['2021-03-17', '1000'],
])('threshold on %s is %s', (date, expected) => { /* ... */ });

I also did mutation checks: break the code on purpose and confirm a test turns red. In one domain, 2 of 7 mutations survived, which showed two blind spots in the tests. I very like this exercise, because it tests the tests.

The test suite grew from 23 tests at the start to around 1,000 at the i18n launch, and today it is over 1,100.

Reconciliation: recompute, don't copy

Migration is where most rewrites cheat a little. I tried not to:

  1. Raw tables are copied (companies, invoices, expenses, assets). Relationships are rebuilt through id_map(table, legacy_id, new_id).
  2. Derived tables are not copied at all (GST returns, income tax results, depreciation snapshots). The new engine recomputes them from the copied raw data.
  3. A script compares every company, every tax year and every GST period, field by field, at cent precision.

Copying only proves that two databases agree. Recompute-then-compare is both the migration and the strongest acceptance test.

Each value gets one label: MATCH, KNOWN(<ledger id>) or UNEXPLAINED. The classifier is code, not my opinion, and a known class must reproduce the exact amount of the difference. For example, for the "unpaid invoice residue" class, the income difference must equal the sum of the stripped rows within one cent.

Result on 1,070 valuesCount
MATCH999
KNOWN71
UNEXPLAINED0

Although the biggest group of known differences looks scary, but it is harmless: 39 values are zero-amount GST periods that the old system created by a bug, and the new system just regenerates them correctly. The full breakdown:

ClassWhat it isValues
D-03Empty GST periods from an old period-generation bug, regenerated39
OPEN-05Income left behind when a paid invoice was set back to unpaid10
D-11Old diminishing-value depreciation used a column nobody updated9
cascadeSummary fields following an already classified component7
REG-03Old summary included assets still pending review3
STALEOld summary row never refreshed after changes2
D-10GST for an unregistered period1

D-11 is the most serious one. The old code calculated diminishing value as baseCost - accumulated_depreciation, but that column was never updated anywhere and was always 0 in production. So from the second year, the old system charged the rate on the original cost instead of on the reducing balance. On the golden sample the old value is 496.98 and the new value is 443.06.

Reconciliation also found bugs in my new code, which was the whole point:

  • Invoice unit prices in the old system exclude GST. My first version treated them as GST-inclusive, even with a comment that said "same as legacy".
  • The expense deduction rate was never applied, so every expense was treated as 100% deductible.
  • The deduction amount must be (total − GST) × rate, and the first version used total × rate.

None of these were found by reading code. They were found by comparing with real, anonymised data. The rule I follow now: when a difference appears, check the divergence ledger first, then suspect the new engine. And any commit that changes a rule compared to the old system must add a ledger row in the same commit.

Data safety: fail closed

Row-level security

CREATE FUNCTION app.current_company_id() RETURNS uuid LANGUAGE sql STABLE AS $$
  SELECT NULLIF(current_setting('app.company_id', true), '')::uuid
$$;

-- created in a loop over every tenant table
USING      (company_id = app.current_company_id())
WITH CHECK (company_id = app.current_company_id())
  • Each request runs in a transaction that sets app.company_id with set_config(..., true), the same as SET LOCAL.
  • Missing context gives zero rows, not an error and not everything, thanks to NULLIF.
  • The app role cannot bypass RLS, and tables use FORCE ROW LEVEL SECURITY, so even the owner is filtered.
  • Policies are created in a loop, because copy-pasted policies drift.
  • RLS is not enough alone: the company ID comes from a cookie, so every request also checks membership. "Missing" and "not a member" return the same error, so IDs cannot be enumerated.
  • Jobs that cross tenants must be declared as system jobs with a written justification.

Money

Money is numeric(18,4) in the database and decimal.js in code. It travels as a string, never a JSON number, and rounding is half-up to match the old PHP round(). An audit still found a bug: an edit form refilled amounts with Number(total).toFixed(2), so 115.1234 silently became 115.12 when you saved without changing anything.

Idempotency and migrations

  • Every write needs an idempotency key. It is required, not optional, because the weak-network mobile retry is exactly the path most likely to forget an optional key.
  • The key row and the business write are in the same transaction. A SHA-256 fingerprint of the procedure plus the canonical input detects the same key reused with a different body.
  • Migrations are expand-and-contract: add, confirm, then delete in a separate step. CI rejects any pull request that edits an existing migration file.
  • Because migrations only expand, old code still runs on the new schema, and that is what make rollback possible.

Email and AI, also designed not to lie

Outbox email. Sending an invoice writes a row into an outbox table inside the same transaction as the status change. A worker job runs every minute, claims rows with FOR UPDATE SKIP LOCKED, and sends outside the transaction. So "status changed but email not queued" and "email sent but transaction rolled back" cannot happen. A send that times out is marked failed and not resent, because a customer receiving the same invoice twice is worse.

AI document inbox. One rule: AI only suggests, it never books.

  1. The model reads a receipt.
  2. Zod validates the response: amounts must be decimal strings, currency must be NZD, an unknown category becomes null.
  3. Deterministic rules in the domain package decide: expense, asset or personal purchase.
  4. The user confirms every posting.

One bug from this area was very surprising. The claim query used WHERE id IN (SELECT ... LIMIT n FOR UPDATE SKIP LOCKED). With RLS, the planner chose a nested loop and ran the subquery again for each outer row, so a batch size of 2 claimed 3 rows. It only reproduced on the shared dev database.

-- simplified shape of the fix
WITH picked AS MATERIALIZED (
  SELECT id FROM outbound_emails
  WHERE status = 'queued'
  ORDER BY created_at
  LIMIT $1
  FOR UPDATE SKIP LOCKED
)
UPDATE outbound_emails SET status = 'sending'
FROM picked WHERE outbound_emails.id = picked.id;

Releases with gates

There were 13 production releases in the first two weeks. Each one follows the same steps:

  1. Preflight on the host: disk capacity must pass.
  2. Image tags set to the commit SHA, never a moving tag.
  3. docker compose up -d
  4. Smoke script: web health with retries, the PDF service, the redirect from the old host, storage returning 403 to anonymous users (which proves the request really reached storage), and the backup service running.
  5. The env file is saved with the previous SHA before every release, so rollback is two image tags and the same smoke test.

Some lessons came from real releases:

  • One release pulled new images, but the old compose file was still on the server, so a new environment variable was missing. Now the release compares checksums of local and remote files first.
  • A sed replacement silently did nothing because the key did not exist yet, and the AI service started disabled. Now environment edits are upserts.

What I would tell myself before starting

  • Write the yardstick before the thing you measure.
  • Recompute instead of copying, and give every difference a name.
  • Make the safe behaviour the default: zero rows when context is missing, required idempotency keys, expand-only migrations.

None of this is fancy. It spend more time at the start, but it is the reason I can say the new numbers are right, and show the evidence.