Our work
We developed the monitoring and emergency-response system for GridLockNode. The work covered transaction simulation, risk evaluation, isolated signing and scoped circuit-breaker contracts across Ethereum, Arbitrum One and OP Mainnet.
An automated security system needs enough authority to interrupt an attack. That authority creates a second problem. If a detector is wrong or its signing path is compromised the defence can become a source of disruption. For GridLockNode we bounded emergency authority in the contract, outside the monitoring software itself.
GridLockNode is an automated smart contract risk monitoring and circuit breaker system for Ethereum, Arbitrum One and OP Mainnet. It observes transaction activity, simulates execution and evaluates risk before submitting a scoped pause to the protected protocol. Our approach gives the automation one privileged action. It can suspend an affected pool, market, vault or registered function. Asset custody, upgrades and recovery remain under separate authority. This boundary shapes the detection pipeline, signing infrastructure and operational response.
Emergency authority had to be narrow enough to trust
A monitoring service needs to process untrusted transactions and react to changing network conditions. Giving that service a general administrative key would connect a large external attack surface to protocol control. We separated the security plane from the application stack and restricted its on-chain role through the EmergencyCircuitBreaker contract. Regional signer accounts receive PAUSER_ROLE for registered scopes. OpenZeppelin role primitives enforce the permission boundary in the contract rather than relying on a daemon to follow an internal policy.
A scope identifies a specific protocol unit or action. Deployment tooling maps protected entrypoints to stable bytes32 identifiers so the coordinator can select the smallest sufficient response. An oracle incident affecting one market can pause that market without automatically disabling unrelated vaults. Ordinary automated signers receive neither the global pause role nor permission to change the scope registry. This makes the extent of an emergency action part of the deployed permission model.
The configured on-chain role excludes the following operations.
- Transfer user funds or protocol treasury assets.
- Upgrade contracts or change oracle and collateral configuration.
- Grant roles or expand its own emergency permissions.
- Unpause a protected scope after an incident.
The restriction continues through the signing path. Detection workers have no KMS access. The incident coordinator creates an immutable PauseIntent, a nonce agent controls transaction sequencing and an isolated signer proxy validates the request before invoking a non-exportable secp256k1 key. That validation covers the chain, destination contract, function selector, scope and fee limits. Services authenticate through workload identities and mTLS. The relayer receives signed transaction bytes; private keys remain in the isolated signing service. A compromised monitoring component can still threaten availability, but its pauser credential does not provide a direct withdrawal or upgrade path.
The response path needed speed without losing evidence
Every dependency between observation and submission consumes part of the response window. We kept durable queues, analytical databases and SIEM integrations outside that path. The regional pipeline uses bounded gRPC calls to move a candidate through simulation, risk evaluation, incident authorisation and signing. Evidence is mirrored asynchronously into an append-only ledger. If the audit sink becomes unavailable the system writes to a local encrypted write-ahead log. Reaching the log's safety watermark raises a degraded-audit condition rather than silently discarding the record.
Simulation capacity also needs protection from the traffic being investigated. The ingress service filters candidates by protected addresses and function selectors before sending them to expensive tracing workers. Direct calls to protected core contracts receive reserved worker capacity. Lower-priority contextual traffic has separate bounded queues and can be displaced under load. Reserved workers keep core-contract tracing separate from lower-priority traffic. Trace execution is bounded by resource limits and timeouts; worker availability still depends on provisioned capacity.
We recorded observation, simulation, authorisation, submission and on-chain confirmation as separate stages. An accepted RPC submission does not establish that a builder included the pause before a harmful transaction. Operators can inspect submission timing and the resulting contract state independently.
Each chain changes what an early response can mean
Ethereum public transaction feeds can expose a candidate before execution. GridLockNode combines locally operated observation with supplementary provider streams, then performs state-critical simulation against locally controlled execution nodes. The relayer uses private-first submission to reduce unnecessary exposure of the defensive transaction and follows a defined fallback policy when that route fails. EIP-1559 fee replacement operates within configured ceilings. Higher fees can improve competitiveness but cannot force a builder to choose a particular ordering.
Arbitrum and OP Mainnet require different assumptions. The system does not depend on access to a public Ethereum-style sequencer mempool on either network. For transactions submitted through a protocol-controlled route it can run preflight before submission. For external activity the first usable evidence may arrive only after sequencing or execution. The response then aims to contain subsequent activity. Chain-specific adapters handle observation and approved submission routes so these differences remain explicit in both execution and reporting.
This boundary keeps synchronous protections inside the protocol. An external monitoring service cannot undo a completed atomic exploit when its first evidence is the executed transaction. Reentrancy guards, oracle deviation checks, stale-price checks and outflow limits therefore remain synchronous controls in the EVM execution path. GridLockNode adds detection and emergency containment around those controls. It does not require operators to assume that every attacker will expose a transaction early enough to intercept.
The system can react quickly because its authority is narrow. It can pause a defined scope while custody, upgrades and recovery remain separate. That boundary lets the protocol automate emergency action without giving the monitoring stack general control over the assets it protects.
Detecting attacks without making disruption cheap
Once a detector can pause a market its false positives become an attack surface. A large flashloan, deep call tree or sudden price move may look suspicious while still being legitimate. If any one of those signals automatically stopped the protocol an attacker could reproduce it to deny service without stealing funds. We made the decision path depend on the economic effect of a transaction and the protocol conditions it violates.
The simulation service runs against locally controlled Reth nodes and pins every evaluation to a concrete block hash. It extracts call trees, state changes, token balance deltas and relevant storage accesses. Recording the block reference prevents two workers from quietly evaluating different versions of a moving latest state. For order-dependent Ethereum candidates the service can evaluate up to three bounded scenarios. These cover the canonical head, a relevant set of pending predecessor transactions and a configured adverse state within valid on-chain limits. The aim is to expose meaningful dependencies without claiming to reconstruct the future block.
The risk engine uses two decision paths. Explicit hard invariants run first. A violation of a configured accounting floor, collateral bound or outflow limit can authorise a scoped pause without waiting for a weighted anomaly score. Each invariant has a reviewable predicate, tolerance, state source and maximum permitted scope. Soft signals pass through a separate correlation policy. Statistical decisions require corroboration through independent detector instances and evidence families or another defined policy condition. A deterministic accounting failure cannot be averaged away by unrelated low-risk features.
Oracle manipulation illustrates the difference. A large DEX trade alone does not justify pausing a lending market. The detector examines price deviation and whether a protected borrow or mint operation actually consumes the distorted value. Flashloan size contributes context only when it correlates with another relevant signal. The decision is therefore tied to a harmful use of the changed state rather than the size or unfamiliarity of the transaction.
Some dangerous dependencies only exist during execution
Final balances do not explain every attack path. In read-only reentrancy one contract may expose a temporary price or reserve value before restoring its accounting invariant. Another contract can read that value during a callback and use it in an economic action. The second contract does not need to be re-entered itself. A detector focused only on repeated function calls could miss the dependency.
GridLockNode traces relevant SSTORE writes and SLOAD reads at execution-frame level. The extractor connects a temporary change in a protected producer contract to a downstream consumer only when control crosses an external boundary and the value affects a protected economic action. Semantic tags identify values such as share price, total assets and collateral value. A normal write followed by a normal read is insufficient evidence. The detector looks for an unrestored producer invariant and a material effect on the consuming protocol operation.
An implementation-diverse shadow backend provides a further check for ambiguous or high-impact evidence. Independent replay uses a different execution implementation from the primary engine. Requiring both backends to finish every evaluation would put client diversity directly into the latency budget. The architecture instead permits a reproducible hard-invariant response to proceed while shadow verification runs in parallel. Agreement checks for ambiguous statistical evidence follow the configured policy and response window. Backend disagreement remains part of the incident record.
Redundancy had to keep dependencies off the incident path
Two active regions improve resilience only if they can complete an emergency action independently. Sharing one signer account would make both regions compete for the same nonce sequence and introduce distributed coordination during an incident. GridLockNode assigns independent pauser identities to each chain and region. A single nonce agent owns each signer account. Either region can submit a pause without waiting for the other to allocate a nonce or acknowledge ownership.
The contract makes repeated pause calls idempotent so concurrent regional responses converge on the same paused state. Transaction replacements reuse the allocated nonce with bounded fee escalation. This separates redundant detection from transaction sequencing and avoids turning a regional outage into a shared signing bottleneck. The trade-off is that each regional identity needs its own role management, monitoring and gas funding.
Funding is kept outside the breaker contract. A pauser needs native tokens before it can submit the defensive transaction, so reimbursement after execution would not solve an empty account. An isolated balance-runway mechanism handles replenishment without giving the pauser access to treasury administration. The breaker itself neither transfers assets nor maintains a refund pool. Its contract surface stays focused on changing pause state and recording the reason.
An accepted transaction is not a contained incident
A relay acknowledgement confirms receipt of a request. It does not establish that the intended scope is paused. The chain watcher closes that gap by checking that the defensive transaction belongs to the current canonical chain, that its receipt succeeded and that isPaused(scope) is true at the observed head. The incident reaches PAUSE_CANONICAL only after all three conditions hold. If a reorganisation removes the pause transaction the watcher reopens the incident and the submission process resumes.
The same rule applies when an incident affects more than one network. A shared bridge adapter or oracle dependency may require action on several chains. GridLockNode sends an authenticated emergency envelope through a fast off-chain channel and retains a separate canonical cross-chain evidence path. Each destination validates the origin signer, policy version, expiry, evidence reference and registered scope mapping before invoking its own local signing pipeline. Duplicate messages converge through an idempotency key and stale messages are rejected before signing.
A source incident cannot automatically expand a market-level pause into a global shutdown on another chain. The destination preserves or reduces the mapped scope unless a separately authorised emergency policy permits broader action. Status is tracked independently for every destination. Ethereum and Arbitrum can be paused while OP Mainnet remains in a retry state. Operators can see that partial result instead of receiving a misleading global success flag.
Recovery deliberately follows a slower path
Fast pausing and fast unpausing create different risks. An automatic return to service could reopen a vulnerable scope before a fix or accounting check is complete. Recovery therefore belongs to a separate authority and follows an enforced two-step process. The recovery authority schedules an unpause, waits for the mandatory delay and then executes it. The interval provides time to revoke a suspect pauser, replay the incident and verify balances or a protocol fix. Even a confirmed false positive follows this recovery path.
The release process applies a similar separation to new detection policies. A policy first runs in shadow mode where it observes transactions, evaluates evidence and constructs hypothetical pause intents without signing them. Review includes false-positive pauses, simulation latency, timeouts and disagreement between backends or regions. Historical fork replay and benign transaction fixtures check both exploit recognition and legitimate behaviour. Controlled emergency drills then exercise real signing, submission, on-chain pause verification and scheduled recovery through a registered non-economic test scope.
Every automated action links back to the triggering transaction, pinned simulation state, extracted features and detector and policy versions. The record continues through signer identity, nonce, relay responses and the canonical pause state. That evidence gives operators a way to investigate why a market stopped and lets reviewers reproduce the decision. A false-positive pause becomes a release-blocking event for the affected detector policy until its cause and regression coverage are established.
The delivered operating model makes those decisions inspectable at the point where they matter. Responders can identify the exact scope that stopped, verify whether the pause remains on-chain and determine which authority can restore service. The incident record separates chain-specific response timing, detector decisions and recovery actions through the same contract boundary.
See our architecture in practice.
DEVLAB · ARCHITECTURE EXAMPLE
Agent
Commerce
A look inside the software architecture behind Agent Commerce.
View architecture