We learned to preview every transformation, scope initial writes to one tenant, load shared lookup data once, preserve separate display and storage identifiers, and reject invalid financial rows before persistence. WebEdify, MACRIM's .NET application platform, supported that approach with markdown migration commands and WebEdify.Billing. In this benefits-administration migration, those commands reshaped nested legacy SQL data into Billing documents and wrote them only when an operator explicitly selected commit=true.

Why was invoicing the hardest part of the migration?

Contacts, invoices, and transactions in the source system combined relational columns with nested JSON. That pattern is convenient until you need a normalized billing UI:

  • Payment methods sat in a nested contact collection — not a separate table.
  • Mailing addresses sat in another nested collection.
  • Invoice line items sat on the invoice header — no standalone line-item table.
  • Ledger rows were close to Billing vocabulary but used source-specific aliases and status conventions.

WebEdify.Billing expects first-class documents: payment_method, contact_address, invoice_item, invoice_transaction, with segment:id keys and normalized dates and totals. The gap is structural, not philosophical. You are not "modernizing" by rewriting screens first; you are reshaping documents so existing admin and client invoice templates can bind to Billing search and ledger commands.

Can you migrate invoicing without a big-bang rewrite?

Yes — as a command chain, not a weekend script dump. The pattern we used:

Step Purpose
Contacts Copy contact attributes; leave nested collections for separate passes
Contact addresses Explode addresses into first-class records
Payment methods Explode tokenized payment references into first-class records; never migrate CVV
Invoices Normalize header dates, totals, customer references, and billing fields
Invoice items Explode nested lines into first-class records
Invoice transactions Remap ledger fields; join payment-method metadata once for display fields
Full chain Run every step, optionally scoped to one tenant segment

Each step is a markdown command with **Roles:** admin, fail-fast error mode, and commit=true as the only gate to writing. Omit it — or set it false — and you get a dry-run: transform only, no Cosmos writes. That is how you validate field maps on one segment before touching production volume.

Forward progress on admin invoicing did not wait for every downstream workflow to finish. Headers, items, and ledger rows in Billing shape were enough to wire list, compose, client pay, and print flows — the same invoicing and receipts surface the marketing site describes.

How do you run a migration safely at scale?

Three rules we would repeat on any similar engagement:

1. Dry-run is the default. Treat every migrator as read-only until commit=true. Review counts and sample documents in the working tree.

2. Scope with segment when learning. Run one plan end-to-end before omitting segment to mean "all." Failures are easier to reason about when the blast radius is one partition.

3. Bulk writes with bounded concurrency. Serverless Cosmos throttles unbounded parallel upserts. We used savebatch with bulk mode and bounded concurrency, tuned to the target environment. Upserts on stable keys are idempotent; a partial run can be re-run without inventing new ids.

Volume caps belong in a dedicated rowlimit parameter on transform SQL, not in a command option whose meaning belongs to the surrounding list operation. Use that explicit parameter in SQL TOP clauses so the migration limit remains unambiguous.

When should you join in memory instead of querying per row?

Two places where naive SQL or cache lookups hurt:

CSV invoice import. A periodic file lists customer keys, invoice numbers, amounts, and dates. Each row must resolve to a tenant segment and primary contact. A per-row lookup approach multiplies round trips at file scale. The fix: list the segment lookup once, keep it in the working tree, and join it into inbound CSV rows on a stable composite customer key. The reusable pattern is one list, one in-memory merge, then shape.

Transaction card display. Stamping card_last4, card_type, and display text on historical payments could be done with per-row SQL against every contact's wallet JSON. That created unnecessary repeated work. Instead: list all wallets once, rename to a sibling node, join into transaction results on 1bf958c6-1fe5-48e3-8075-1b52ee906969|, then save.

The lesson is not "joins are fast." It is: know when your enrichment is O(rows × lookups) and collapse lookups to O(1) lists before you iterate.

What identity rules prevent a bad file from overwriting the wrong invoice?

Display identifiers and storage identifiers must diverge for financial documents.

Legacy SQL → Cosmos: Preserve the source catalog as an attribute and map the stable catalog to id for the document key. Partition as segment:id. If another source id differs from that catalog, preserve it separately — do not silently merge two keys.

CSV import: Human invoice_number stays on the document for search, email, and print. id / catalog are opaque: derived from encrypted material over invoice number plus normalized issue date, URL-safe. Re-import of the same number and date upserts the same document; a typo in the number does not clobber an unrelated row. issue_date is required; fail the entire file if it is missing or unparseable unless you explicitly set force=true, which moves bad rows to an unmatched bucket and continues with matched rows only.

Forward-only note: rows imported earlier under raw invoice numbers as ids do not automatically merge with opaque-catalog imports. Plan identity cutovers explicitly.

Short declarative sketch of the import join (illustrative):

**command** models/tenant/list-import-lookup
- onto: $
- as: segmentlist

**read** `` as csv using invoice-import-map
- onto: $
- as: inbound

**join** $.segmentlist into $.inbound match |
- including: catalog,primary,email,name
- prefix: seg_

Validation and opaque id shaping follow in a small code step where encryption service access is required — the only place code was justified; explode and join stayed declarative.

What broke in practice — and what we did about it?

Issue Fix
Cosmos rejected "unset" partition keys Use segment:id with id populated from catalog
Nested-path elimination targeted the wrong level Eliminate nested collections relative to each contact, not from the document root
A list option and SQL row cap used the same name Use @rowlimit in transform SQL
SQL CONCAT collation errors COLLATE DATABASE_DEFAULT on catalog fragments
Historical payments reference missing payment-method keys Leave display fields blank and preserve the source reference rather than inventing data

When the source contains an orphan payment-method reference, migration preserves it; the Billing UI can leave card display text empty rather than inventing a value.

Does this mean every migration is configuration-only?

Invoicing sat in the sweet spot: repeated list → merge → join → savebatch patterns, plus two compact transforms for header normalization and import validation. If a legacy shape requires a large imperative block, treat that as a missing reusable capability rather than burying the migration in a long tenant-specific code step.

For a broader modernization frame — one screen at a time without a rewrite — the same platform supports standing up admin invoicing on Billing models while legacy SQL remains read-only for the next migration pass. You do not need a feature freeze on the old system to prove the new one.

If you are weighing a similar move, start with one segment dry-run of headers and ledger rows, then compare admin search and client pay against legacy. Talk to us with the shape of your legacy data and the workflow you need to preserve; we can determine whether it fits the same command pattern or needs scoped implementation from the pricing page.