Back to selected work

AspanDigital. Keeping payments moving while the system changes

Dmitry Sergeev·13 min read

Category:Payments

Project status:Discontinued by client

Tags:
  • Cross-border payments
  • Stablecoins
  • Compliance
  • MPC custody
  • Double-entry ledger
  • Zero-downtime releases

Our work

We developed the payment, ledger and custody integrations for AspanDigital, including compliance checks inside the MPC signing path. We also built the release and recovery workflows around in-flight transfers, reserved liquidity and versioned payment execution.

AspanDigital is a cross-border payment platform designed around Kazakhstan's trade corridors. Customers receive a final price before authorising a transfer, compliance approval is required inside the signing infrastructure, and customer liabilities are reconciled against custody and blockchain balances. Those promises have to survive software releases, sanctions updates and transfers that remain under review for days. The architecture therefore treats deployment behaviour as part of payment correctness.

The design starts with three boundaries. PostgreSQL's double-entry ledger records customer money. Compliance decides whether a specific movement is permitted. Custody signs only what that decision authorises. We connected those boundaries through payment records, liquidity reservations and versioned signing approvals. The release process preserves in-flight payments while replacement services take over.

Compliance approval belongs inside the signing boundary

An approval flag in an application database leaves the payment service responsible for checking it. A compromised orchestrator could skip that check or change the destination afterward. AspanDigital uses a signed Compliance Attestation that every participating MPC node verifies independently. Without a valid attestation, an honest node refuses to join the signing round. The 3-of-5 quorum places enforcement across separate signing domains rather than leaving it with one application process.

The attestation binds the decision to a canonical TransferIntent containing the chain, asset, source, destination, amount in base units and transfer identifier. Each signer decodes the unsigned transaction itself and reconstructs that intent. It compares the resulting hash with the approved hash instead of trusting a description supplied by the orchestrator. Minimal chain-specific decoders accept only supported transaction shapes. A substituted token, recipient or amount fails the comparison, and an unsupported call structure is rejected.

Before contributing its signature share, each node checks the complete authorisation context.

  • The attestation has a valid signature from the Compliance Engine's HSM-backed key.
  • The decoded transaction matches the approved intent.
  • The approval is unexpired, persisted and unused.
  • Its sanctions-list versions meet the current enforced minimum.
  • A Ledger-signed receipt proves a matching hold exists.
  • The request stays within the node's local wallet-tier velocity limits.

Treasury transfers pass through the same signing checks. The compliance service cannot substitute for the ledger hold, and the orchestrator cannot substitute its own transaction interpretation for the signer's decoder. This closes the application-level bypass under the stated quorum trust model. It does not make upstream screening infallible or guarantee that an external institution will accept the payment. The enforceable property is that the signing path requires the recorded approval for the actual transaction.

Signing checks the current approval version

A sanctions addition can arrive after screening but before signature generation. AspanDigital closes that gap with monotonically versioned lists and a signed MinListVersion record. When an update adds designations, signer nodes reject approvals below the new minimum even if their normal expiry has not elapsed. The orchestrator requests a fresh decision and continues only if the transfer remains allowed. Removal-only updates do not raise that minimum because they cannot turn an earlier allow decision into a block.

We tied sanctions-list activation to the screening fleet and signing nodes. The new minimum version is published after serving screening instances load the corresponding index; lagging instances leave rotation. Signers receive pushed updates with a polling fallback, and approvals based on an older list return for fresh screening before signing.

Inbound funds follow a different path because the platform cannot prevent someone from sending to a public address. Each customer has a dedicated deposit address per chain. Deposits remain there during source-risk screening and Travel Rule matching, and only cleared funds are credited and swept into omnibus treasury wallets. Flagged deposits enter quarantine and case management. The extra sweep costs a transaction, but it prevents automatic commingling of a flagged deposit with treasury funds.

Screening is adapted to the names encountered in these corridors. The Rust engine generates Kazakh romanisation variants, Russian transliterations, Uzbek Cyrillic and Latin forms, and Chinese Han-to-pinyin forms with alternate name orderings. Patronymics receive their own treatment, while phonetic keys and character n-grams retrieve candidates. Dates of birth and strong identifiers help distinguish similar names. Screening-engine review treats recall on the labelled true-match set separately from precision. A smaller review queue does not by itself establish better screening.

A 30-second fixed quote is backed by reserved liquidity

The customer-facing quote includes the source amount, destination amount and all fees, fixed for 30 seconds. Supporting that commitment requires more than caching an exchange rate. The Corridor Router models currencies, chains and payout partners as a graph of value locations and conversion or settlement operations. It filters routes by compliance eligibility, available liquidity, operating windows and amount limits, then ranks eligible paths using fees, spread, latency, reliability and liquidity utilization. The cheapest displayed rail is not necessarily the lowest-cost permitted route that can complete the payment.

When a quote is issued, the ledger reserves funds in every pre-funded pool needed by the selected route. Treasury records the FX exposure and hedges when the net position exceeds its internalisation limit. Smaller opposing flows can offset each other. The stored quote carries its exact route, prices, reservations and pricing-engine version. Acceptance executes that record without recalculating it, including when a newer pricing release is already running. Expiry releases the reservations through ledger-owned hold timeouts.

This moves market and network-fee variation within the quoted promise onto the platform, so quoting also needs exposure caps and loss budgets. Stale prices stop new quotes for the affected pair, and unhealthy partner edges leave the routing graph. Pre-funding makes delivery possible without waiting for a fresh treasury transfer on every payment. Transfers between AspanDigital customers go further: they settle entirely in the ledger, with no blockchain transaction or network fee, while still passing applicable compliance checks. Partner flows can be netted before external settlement.

Reconciliation and verification answer separate questions

Live Solvency Proofs combines continuous reconciliation with published reserve and liability snapshots. On each new block, the system compares ledger reserve accounts, addresses controlled by custody and balances observed by its own chain indexers. Comparisons use the same block cursor and account for pending inbound transfers, pending outbound transfers and fees in transit. A residual beyond dust tolerance opens a classified reconciliation break. An outflow with no matching attestation is a critical incident and stops signing for the affected wallet.

Customer-verifiable snapshots are built daily and on demand for auditors. A Merkle-sum tree commits to customer liabilities per asset, while signed address-ownership messages and block references identify the reserves being compared. Each customer can download an inclusion path and verify that their balance appears in the published liabilities root using an open-source verifier. Snapshot-specific identifiers, salts and shuffled leaves reduce linkage, although sibling sums remain visible in this construction. Continuous reconciliation detects operational divergence; the inclusion proof lets a customer check their place in a published snapshot. Neither establishes that every possible off-ledger obligation has been disclosed.

Deployment starts by reserving capacity for both versions

The no-maintenance-window requirement led to a stricter release objective: keep serving capacity intact and constrain latency changes throughout the rollout. Kubernetes replacing healthy pods successfully is insufficient if new pods saturate storage, connections or CPU while starting. The architecture reserves surge capacity in addition to site-failure headroom. Low-priority placeholder pods hold that capacity until the scheduler replaces them with real surge pods, avoiding dependence on new bare-metal capacity arriving during a release.

Images are pre-pulled onto target nodes from site-local Harbor mirrors before traffic moves. Critical workloads use Guaranteed QoS, with pinned cores for latency-sensitive services and placement rules that spread warming pods across nodes. PostgreSQL connection budgets include both normal replicas and surge replicas. PgBouncer enforces those limits, and connection pools grow gradually with jitter. HSM sessions use a shared per-site proxy, so adding application pods does not multiply sessions against the hardware.

New pods load verified policy bundles, warm route and registry caches, establish connections and run side-effect-free self-tests before becoming ready. Envoy then increases their traffic share over a 60-120 second slow-start window. Retiring pods first stop advertising readiness but keep serving while endpoint removal propagates, then use HTTP/2 GOAWAY to redirect new streams. Existing requests finish within their deadlines, workers stop claiming new batches and consumers commit completed work. Startup and shutdown are separate protocols because each can disrupt payments in a different way.

Canary analysis compares equally fresh processes

Comparing a cold canary with stable pods that have accumulated days of cache state would confuse startup effects with code regressions. Argo Rollouts therefore starts fresh baseline pods of the old version alongside the canary. Both receive comparable traffic in the same window. Synthetic transfers reach the canary before customer traffic, followed by progressive traffic steps. Minimum sample requirements prevent a quiet period from being mistaken for a successful experiment.

Analysis combines a one-sided Mann-Whitney U test at a 0.01 significance level with an explicit p99 regression budget of at most 5%. It also examines errors, CPU per request, memory growth, query fingerprints, downstream failures and consumer lag. Compliance releases compare decision distributions so an unintended change in allow, review or block behaviour is not hidden by good HTTP latency. Failure returns traffic to the stable version, drains the canary and blocks promotion. The stable fleet retains capacity for that abort, so rollback does not depend on provisioning replacement resources.

Promotion moves through the synthetic cell, a small customer cell, the first Almaty site and then the remaining sites. Astana follows as the disaster-recovery standby. Each wave requires analysis and corridor probes, while site headroom allows the affected site to be drained. A controller-enforced site change lock prevents an application rollout from overlapping another disruptive maintenance operation.

Bounded migrations protect the commit path

The ledger's synchronous replication couples background writes to customer commit latency. A backfill can touch unrelated rows yet generate enough WAL to slow the standby acknowledgements that customer transactions await. Isolating its worker pods does not remove that database dependency. AspanDigital therefore treats background migration as a budgeted workload measured against the commit path itself.

Schema evolution follows expand, dual-write, backfill, switch reads and contract across compatible releases. Indexes use concurrent construction where supported, constraints are introduced separately from validation, and type changes use replacement columns rather than in-place table rewrites. Migration sessions have a two-second lock timeout and bounded retries. Even a brief exclusive lock request can queue behind a long query and block subsequent traffic while waiting, so the bounded timeout stops the migration attempt instead of leaving it in that queue. Jobs run once per cluster under an advisory lock, never on application startup.

We kept backfills in short, checkpointed transactions and tied their pace to the live database. After each batch, the controller checks WAL generation, standby lag, customer commit latency, lock waits and table bloat. Crossing a configured budget shrinks the batch and backs off; persistent breaches pause the job. Query-plan checks and canary analysis by query fingerprint address regressions that schema compatibility alone cannot catch.

Stateful services need handover, not just restarts

Kafka consumers use KIP-848 where their clients support it, with cooperative-sticky assignment as the fallback. New consumers join before old consumers drain, allowing partitions to move incrementally. Members retaining their assignments keep processing instead of joining a group-wide stop. A transferred partition can still pause during handover, so consumer lag remains a release metric. Static membership limits unnecessary movement for consumers with substantial local state.

Nonce allocators and outbox relays require exactly one active writer per key. Their successors load current state and shadow the incumbent before leadership changes. A PostgreSQL lease carries a monotonically increasing fencing token, and downstream components reject work bearing an obsolete token. During a wallet handover, allocation requests wait briefly in a bounded queue while the successor takes authority. Shadow divergence aborts the handover before it begins. The fencing check prevents a delayed old process from continuing to act after its replacement becomes leader.

Flink monitoring jobs use a parallel transition because stopping a job to restore windowed state would delay alerts. The old job continues after producing a savepoint. The new version restores onto separate TaskManagers, catches up and writes to a shadow topic. Output parity is checked before case management switches at an aligned offset boundary, with deterministic alert IDs deduplicating overlap. Stable operator UIDs and schema-aware serializers preserve the mapping between saved state and the new topology.

In-flight payments retain their original logic

A transfer can outlive several application releases. Each therefore stores a flow_version, and the orchestrator dispatches through handlers keyed by version and state. New releases can introduce new flows while retaining handlers for every version with unfinished work. The release gate checks that live set and rejects a build that removes a required handler. An exceptional migration uses an explicit, tested upcaster rather than silently interpreting old state under new logic.

The same compatibility discipline covers commercial and cryptographic objects. Quotes retain stored routes and prices, ledger holds retain their own expiry, and durable timer rows survive process replacement. Signers accept a new attestation format before the compliance service begins issuing it. MPC presignatures are tagged by protocol version, with the new pool filled before the old version is retired. These ordering rules let mixed versions coexist without repricing a payment or leaving signing waiting for preprocessing.

Policy and sanctions updates without restarting readers

Policy bundles are signed, immutable artifacts compiled to Wasm. Candidate policies are evaluated against archived decisions and then through asynchronous shadow workers using frozen live inputs. A bounded queue separates that work from the live decision path. If shadow capacity is exhausted, samples are dropped rather than delaying payments. Activation records the effective time and bundle version, and changes requiring reevaluation send affected unfinished transfers through a fresh decision.

Sanctions indexes use a different mechanism because readers need a complete, internally consistent dataset. Rust workers build and test the next index on reserved cores while the current index serves requests. Memory is sized for both. ArcSwap publishes the new index atomically, and existing readers finish against their previous snapshot. Results carry the exact list version used. That version then connects back to the signer-side minimum, allowing hot updates without accepting partially built data or preserving stale approvals indefinitely.

MPC upgrades preserve a local signing quorum

Signer rollout follows the custody topology. Four of the five nodes sit in distinct Almaty-region security domains, with the fifth in Astana. Taking one local node out still leaves three local nodes for the online quorum. The coordinator stops assigning it new sessions, drains active sessions and excludes presignatures involving that node. Reserved presignature capacity covers the remaining quorum, so a planned upgrade need not move online signing to a higher-latency regional path.

Node replacement uses the reviewed build digest, and security officers add its confidential-VM measurement to the peer allowlist before it joins. The node resumes preprocessing, passes synthetic signing checks and bakes for at least 24 hours before the next node changes. New protocol features activate only after all five nodes support them. Share refresh, key rotation, HSM firmware and removal of the old measurement are separate changes. Keeping an online quorum in the same region avoids routing ordinary signing through the disaster-recovery site during a node replacement.

The release design addresses each source of interference with a specific control.

RiskEffectControl
Surge demandResource contentionReserved capacity
Image pullsI/O pressureLocal pre-pull
Cold podsSlow requestsWarm-up and ramp
Pod shutdownDropped requestsDrain and GOAWAY
DDL waitsQuery queuesLock timeout
Backfill WALSlow commitsLag-based throttle
RebalanceConsumer pausesKIP-848 fallback
Writer restartConflicting authorityLease and fencing
Flink restoreDelayed alertsParallel jobs
List rebuildScreening stallsAtomic swap
Flow changesBroken transfersVersioned handlers
Signer restartQuorum latencyLocal reserve
Bad releaseWider impactCells and abort

AspanDigital treats the right to move money, the price already promised and the resources needed to finish a payment as commitments that survive a release. Attestations, ledger reservations and versioned execution protect those commitments, while isolated capacity and measured promotion keep deployment work from competing with customer payments.

See our architecture in practice.

DEVLAB · ARCHITECTURE EXAMPLE

Agent
Commerce

A look inside the software architecture behind Agent Commerce.

View architecture
Agent Commerce — Software Architecture, designed by Dmitry Sergeev