Most online converter ask you to upload your file, and you never really know what happens after. I wanted to build tools that tells you the truth on every page: this file stays in your browser, or this file goes to our server, and here is what we do with it.
The result is Kivoza:
- image converters: HEIC, WebP and AVIF to JPG or PNG
- an OCR tool for printed text
- five PDF tools: merge, split, rotate, extract pages, images to PDF
- a Markdown site. There have eight Markdown converters: PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON and EPUB.
One repo, seven small sites
| Site | Build | Where files are processed |
|---|---|---|
kivoza.top | Plain HTML + esbuild | No tools, guides only |
heic. / webp. / avif. | Vite, about 10-line config each | Browser only |
ocr. | Vite + asset copy plugin | Browser only |
pdf. | Vite, multi-page | Browser only |
markdown. | Vite, multi-page | Browser for 4 formats, server for 4 formats |
They are all built from one repository with one lockfile. Each site have its own Vite config, and for the simple ones it is almost nothing:
export default defineConfig({
root: resolve(process.cwd(), 'apps/heic'),
build: { outDir: resolve(process.cwd(), 'dist/heic'), emptyOutDir: true, target: 'es2022', sourcemap: true },
});
Why not one big app?
- Security headers per site. This is the most important reason. The image, OCR and PDF sites use
connect-src 'self', so the page cannot send your file anywhere even if there is a bug. Only the Markdown site adds the API origin. With one origin, every tool would need the widest policy. - Independent deploy and rollback. Each site is its own Cloudflare Pages project.
- Heavy files stay where they belong. The OCR build is 74 MB because of language models, the PDF build is 10 MB, and neither is inside another tool's bundle.
- Phased releases. No fake entries for tools that do not exist yet, and each tool page works as a direct landing page.
I tried hard not to build a "platform": no shared conversion engine, no plugin registry. The first released tool became a written visual contract (header, hero, converter card, info, limits, FAQ, footer), and each tool copies this structure. Some code is duplicated, and I accept it. Only the Markdown site has a real shared module, and its comment says it never branches on a file format.
What runs in the browser
HEIC: building libheif myself
HEIC is the hardest format, because browsers other than Safari cannot decode it. I first used a prebuilt npm package, but an audit found that it embedded an old version of libheif, older than the version the upstream release notes asked for. So I decided build libheif 1.23.4 myself to WebAssembly with Emscripten, inside a pinned Docker image.
There was one more problem. Emscripten's Embind generates JavaScript at runtime by default, which needs unsafe-eval in the CSP, and that is not safety for a page that handles user files:
# build flag
-sDYNAMIC_EXECUTION=0
# resulting CSP only needs
script-src 'self' 'wasm-unsafe-eval'
Because libheif is LGPL, the site also ships the full source and a rebuild script, so people can check it by themself.
WebP and AVIF: native decode, strict parsers
Browsers decode these natively with createImageBitmap, but I wrote small parsers anyway:
- An animated WebP or AVIF is rejected with a clear message, instead of silently converting only the first frame.
- The pixel limit uses the largest image size declared in the file, so a small thumbnail cannot hide a huge main image.
- For JPG output, transparency is flattened onto white, and the output Blob type, magic bytes and file extension must agree.
Limits, and where they came from
| Limit | Value |
|---|---|
| Files per batch | 8 |
| Size per file / per batch | 20 MiB / 64 MiB |
| Pixels per image | 40 MP |
| Parallel conversions | 2 |
These numbers came from real tests. A batch of eight images near 40 MP once hit the 2 GiB memory limit of my test container. After three changes, the same batch peaked at about 1.64 GB:
- one output canvas, cleared after each
toBlob - previews capped at 1024 pixels
- every decoded libheif image freed in a
finallyblock
OCR and PDF
- OCR uses tesseract.js 7 in a Web Worker. Models are served from the same origin, not a CDN, and cached in IndexedDB. English is about 11 MB, simplified Chinese about 20 MB.
- One nasty bug: if a model download returned HTTP 503,
createWorkerstayed pending forever. The fix tracks native workers, adds a 30-second deadline, terminates orphans and shows a Retry button. - PDF uses a maintained fork of pdf-lib for editing and pdf.js for thumbnails. I rejected MuPDF, Poppler, Ghostscript and PDFium WASM because of GPL or commercial licences and size. Encrypted PDFs are rejected, not decrypted.
Markdown: local where possible, server only where needed
| Format | Where | Engine and notes |
|---|---|---|
| Browser worker | WASM PDF inspector; scanned pages are flagged, not OCR'd | |
| HTML | Browser worker | turndown + detached DOM; parsed into an inert <template>, scripts and on* removed, remote images never fetched |
| CSV | Browser worker | Papa Parse in strict mode; unterminated quotes are fatal |
| JSON | Browser worker | Big-integer parser, duplicate keys rejected, depth up to 100 |
| DOCX, PPTX, XLSX, EPUB | Server | officeparser, exceljs (reads cached formula results, never runs formulas), SSF for Excel formats |
The JSON choice matters: a normal parser turns 900719925474099312345 into a rounded float. My first JSON version used another library that silently dropped keys named __proto__ and constructor, which an audit caught.
Only four format need the server. I evaluated a popular native library first. It passed 31 of 34 checks, but it rounded big integers, had no option for slide notes or sheet selection, and put EPUB chapters in the wrong order.
The upload is honest
- The four server pages show a consent checkbox that is unchecked by default.
- The client checks it again before every single request.
- The page says: the file is sent to the Kivoza conversion service, temporary input is deleted after success, failure or cancellation, results stay in memory for up to 10 minutes, and logs contain no document body.
- The local pages say the opposite, and an audit checked both claims with network capture.
The result is always an editable textarea with a live preview. The preview pipeline is marked, then DOMPurify with a small allow-list (no images, iframes, forms or styles), then a link filter that drops javascript: and data: links. Cancelling a conversion terminates the worker, so the parser really stops.
Hardening the conversion API
ZIP bombs and bad archives
Office files are ZIP archives, and ZIP archives can be bombs. Before officeparser sees any bytes, a preflight streams the archive through fflate in 64 KiB chunks and counts the bytes actually inflated:
maxInputBytes: 50 MiB
maxUncompressedBytes: 64 MiB // counted while inflating, not trusted from headers
maxZipEntries: 512
maxCompressionRatio: 1_000
maxMarkdownBytes: 8 MiB // UTF-8 bytes
maxImagePixels: 40_000_000
It also rejects encrypted entries, unsafe paths like ../, macros (vbaProject.bin), external links and DRM-protected EPUBs. Embedded images are exported only as PNG, JPEG or GIF, and their dimensions are read from the file header before the pixel check.
Process and container
- Each conversion runs in a forked child process with a minimal environment and a 60-second hard kill.
- One active worker and four queued slots; after that the API returns 429.
- The container has no network (
network_mode: none) and listens only on a Unix socket. - Non-root user, read-only filesystem, all capabilities dropped, 2 GiB memory and 512 PIDs.
- An nginx container in front allows one upload per second and only the exact Markdown site origin.
Two bugs that taught me the most
- Characters are not bytes. The 8 MiB output cap first counted characters, so three million Chinese characters (9,000,228 bytes in UTF-8) passed. Now it uses
Buffer.byteLength. - A token in the URL. In the first design the task token was part of the URL path. When I read the source code of cloudflared, I find that it logs the full request URL when dispatch to the origin fails, so a token could leak into tunnel logs. The fix:
The old image is banned as a rollback target.before: GET /tasks/<token> after: GET /api/markdown/task Authorization: Bearer <token> (any query string on this path → 404)
Deploying seven sites without breaking the home page
- GitHub Actions plans which sites changed by path:
apps/heic/rebuilds only HEIC, while a lockfile change rebuilds everything. - On a push to main it compares against the last successful production run, not the previous commit, so a failed run cannot hide earlier changes.
- After each upload, a gate checks the home page, one asset, the security headers and a real 404. That last check exists because Cloudflare Pages served the SPA page with status 200 for unknown paths until each site got a real
404.html. - For the Markdown site, the gate sends a valid-looking but unknown token and expects 404, not 401. That proves the edge really passes the
Authorizationheader through. - The root site is always published last, after all six tool sites are healthy.
One more small thing: Cloudflare automatically injected its analytics beacon. It was blocked by the CSP and also contradicted the "no analytics" text on the pages, so I disabled it per host instead of loosening the CSP.
Tradeoffs I wrote down
- Canvas re-encoding does not keep EXIF, ICC profiles or HDR metadata.
- OCR only reads printed text, in three language choices.
- Four Markdown formats need an upload.
- One worker and a small queue trade speed for predictable memory, which is an useful rule for a free tool on an existing server.
- Most testing was in Chromium, with some Firefox, not on real iPhones.
The main lesson is simple. Decide the privacy boundary first, write it on the page, and then let the architecture follow it: separate sites, strict CSP, browser workers for local work, and one small, locked room for the formats that really need a server.