

Substreams Extended Blocks Data: Data Your RPC Provider Can't Give You
Most applications run on a thin slice of what a blockchain records. A team queries an RPC endpoint, pulls blocks, transactions, receipts, and event logs, and builds on that. It works until the moment the data you need was never in a log. A router that moves funds through an internal call, a contract that changes state without emitting an event, a precise native-balance history: none of that shows up cleanly in a standard RPC response. To get it, developers reach for trace endpoints that are slow, metered, and often switched off.
Substreams take a different route. Firehose extracts every block once and stores it as flat files that Substreams reads in parallel, so the data arrives without the rate limits and repeated polling of live RPC. On EVM chains, that extraction comes in two shapes: the base block model and the extended block model. Which one a chain has decides how much of the chain you can work with. This post explains the difference, and why extended support gives you data that the RPC most developers use today cannot.
Two Ways Chains Get Substreams Support
A chain can be onboarded to Substreams through two integration paths, and each produces a different block model.
A base block comes from an RPC Poller integration. Substreams polls a standard RPC endpoint and captures what that endpoint exposes: blocks, transactions, receipts, and logs. Because it needs only a normal full node for live data and an archive node for history, a base integration is fast to stand up. A new chain can have Substreams support without shipping custom node software.
An extended block comes from full node instrumentation. Instead of polling an interface, an instrumented node emits the complete internal record of each block as it executes: balance changes, internal calls, storage changes, and the rest of the execution detail. This path takes more work to set up, and it produces the richest data model The Graph can serve for an EVM chain.
The same Protobuf structure carries both. Each block records its own detail level, so a Substreams module can tell whether it is reading base or extended data. The practical effect for a developer is simple. On a base block, the fields that only instrumentation can produce come back empty. As the puts it, "The data missing in the Base Block makes the corresponding Protobuf field empty. For example, if you try to read internal calls on a Base Block, the list will be empty."
What Lives in an Extended Block
A few categories of data separate an extended block from a base block. Each one answers a question that logs and receipts alone cannot.
Internal calls: When a contract calls another contract, that inner call is not a top-level transaction and does not appear in the transaction list. A DeFi router that splits a trade across pools, a multisend that pays a hundred addresses, a proxy that forwards every call to an implementation: the value moves in these internal calls. An extended block gives you the full call tree for every transaction. A base block gives you the outer transaction and its logs, and nothing beneath.
Balance changes: An extended block records every change to a native-token balance with the reason for each change. You can reconstruct an exact balance history for any address without re-deriving it from transactions and gas math, and without missing the balance effects of internal calls.
Storage changes: Contract state that never emits an event is invisible to log-based indexing. Plenty of contracts update storage without a matching event, either by design or by omission. An extended block captures the storage slots that changed in each block, so you can index and verify state directly rather than inferring it. It also records the Keccak preimages behind those slots, so a hashed storage key can be walked back to its original mapping key and slot when you know the contract's storage layout.
Code changes: When a contract is deployed or removed, the account's code changes. An extended block records these code changes, so you can track every contract deployment directly, including the contracts a factory deploys through internal calls (CREATE and CREATE2) that never appear as a top-level transaction. Read alongside the full call tree, this lets you follow deployed contracts across the whole transaction.
An extended block also records nonce changes for each account. A base block still carries the data most event-based indexing depends on: transactions, receipts, and logs are all there. What it does not carry is the state-level detail above.
Why Extended Blocks Beat Your RPC Provider
The base model already fixes the delivery problems of live RPC. Firehose extracts each block a single time, so Substreams replays history at high speed instead of hammering an endpoint request by request. There are no per-call rate limits, no missed events during reconnects, and the same input always produces the same output. For a team that today stitches together eth_getLogs and eth_getBlockByNumber against a metered provider, even a base-block Substream is a step up in speed and reliability.
Extended blocks add the part RPC providers make hardest to get. Internal calls, balance changes, storage diffs, and contract code changes are not in the standard JSON-RPC surface. To pull them, you have to call trace methods such as debug_traceTransaction or the trace_ namespace, or run state-diff tracers across archive state. Those endpoints are expensive, slow, and heavily rate-limited. Many providers gate them behind premium tiers or disable them entirely, and not every node exposes them. Building a reliable pipeline on top of them is a project in itself.
An extended block has already captured that data at execution time, from an instrumented node, for every block. Substreams then serve it with the same parallel replay, determinism, and reorg awareness as everything else. The data that would cost you a premium trace plan and a fragile scraping layer is simply a field on the block.
What Extended Blocks Support Lets Developers Build
The three data types map directly to things you can build on an extended chain that are impractical on base blocks or raw RPC.
Full call-tree data makes contract accounting correct. If you index a router, an aggregator, or any contract that moves value through internal calls, extended blocks let you attribute every transfer instead of guessing from top-level transactions. The same data feeds MEV analysis, trace-level debugging, and any product that needs to see what a transaction did, not just that it happened.
Balance-change data makes native-token history exact. Wallets, portfolio trackers, accounting tools, and tax reporting can read balances at any block without reconstructing them, and without the gaps that internal transfers leave in a transaction-only view.
Storage-change data makes event-poor contracts indexable. When a protocol changes state without emitting the event you need, extended blocks let you read the state change straight from storage. You can also verify that indexed state matches on-chain state, which matters for compliance-grade and audit-grade data.
Code-change data makes contract deployment traceable. If you track factories or need a live registry of deployed contracts, extended blocks surface every deployment, including the ones created through internal calls that a top-level transaction list never shows.
When Base Blocks are Enough
Extended is not a strict upgrade for every job. A large share of Subgraphs and Substreams only need events, transactions, and receipts, and a base block delivers all of those with the same speed and reliability advantages over RPC. If your indexing logic keys off logs, a base chain serves you well, and the base model is the reason so many chains have Substreams support at all. The base integration path lowers the bar to onboard a network, so coverage grows faster across the many EVM chains launching today.
The decision comes down to the data your use case touches. If you need internal calls, native-balance history, or storage-level state, you need a chain on the extended model. If you index events, base is enough.
How to Check a Chain's Support Model
page, powered by the ecosystem’s networks registry, is the canonical source for which model a chain has. Every EVM network entry carries an evmExtendedModel flag and a blockFeatures field that reads base, extended, or extended@<block>.
As of September 2026, 17 EVM mainnets and 11 EVM testnets carry the extended block model:
| Network | Network ID | The Graph Market Support |
|---|---|---|
| Arbitrum One Mainnet | eip155:42161 | StreamingFast, Pinax |
| Arbitrum Sepolia Testnet | eip155:421614 | Pinax |
| Arc Mainnet | eip155:5042 | Pinax |
| Base Mainnet | eip155:8453 | StreamingFast, Pinax |
| Base Sepolia Testnet | eip155:84532 | Pinax |
| BNB Smart Chain Mainnet | eip155:56 | StreamingFast, Pinax |
| BNB Smart Chain Chapel Testnet | eip155:97 | Pinax |
| Ethereum Mainnet | eip155:1 | StreamingFast, Pinax, DataNexus |
| Ethereum Hoodi Testnet | eip155:560048 | StreamingFast, Pinax |
| Ethereum Sepolia Testnet | eip155:11155111 | StreamingFast, Pinax |
| HyperEVM Mainnet | eip155:999 | Pinax |
| InjectiveEVM Mainnet | eip155:1776 | StreamingFast |
| InjectiveEVM Testnet | eip155:1439 | StreamingFast |
| Ink Mainnet | eip155:57073 | Pinax |
| Linea Mainnet | eip155:59144 | Pinax |
| Linea Sepolia Testnet | eip155:59141 | Pinax |
| Monad Mainnet | eip155:143 | StreamingFast |
| Optimism (OP Mainnet) | eip155:10 | StreamingFast, Pinax, DataNexus |
| OP Sepolia Testnet | eip155:11155420 | Pinax |
| Polygon Mainnet | eip155:137 | StreamingFast, Pinax |
| Polygon Amoy Testnet | eip155:80002 | Pinax |
| Robinhood Chain Mainnet | eip155:4663 | StreamingFast, Pinax, DataNexus |
| Soneium Mainnet | eip155:1868 | Pinax, DataNexus |
| Soneium Minato Testnet | eip155:1946 | Pinax |
| Unichain Mainnet | eip155:130 | StreamingFast, Pinax, DataNexus |
| Unichain Sepolia Testnet | eip155:1301 | Pinax |
| World Chain Mainnet | eip155:480 | StreamingFast |
| Zora Mainnet | eip155:7777777 | Pinax, DataNexus |
Two of these, Arbitrum One and OP Mainnet, are marked hybrid: they provide extended data from a specific block height onward, and the registry lists the exact cutover so you know where the richer fields begin.
The base model covers a larger and faster-growing set of more chains – among them Blast, Celo, Fantom, Gnosis, MegaETH, and Tempo. These counts move as chains are added, so check the registry for the current state.
Non-EVM chains sit outside this base-versus-extended split. Networks like Solana, NEAR, and Bitcoin have their own native block models in Substreams, each shaped to what the chain records.
Core Takeaway
Base blocks give you a faster, more reliable version of the data you already pull from RPC. Extended blocks give you data RPC keeps behind trace endpoints or never exposes at all: the internal calls, balance changes, storage diffs, and code changes that let you account for value correctly, track balances exactly, index state that no event reports, and follow every contract deployment. Before you build on a chain, check its model in the registry, and match it to the data your product needs. If you don’t see your chain, to talk about an integration or request the chain you need.
About The Graph
The Graph is a suite of blockchain data infrastructure products that extract, process, and deliver scalable blockchain data solutions across 60+ networks. The Graph enables application developers, data analysts, AI agents, and enterprise teams that need structured, real-time access to blockchain data. Products include Subgraphs, Firehose, Substreams, and Amp. As of early 2026, The Graph has served over 1.27 trillion queries to more than 75,000 projects, powered by a network of independent Indexers around the world.
Follow The Graph on , , , and . Join the community on The Graph’s , join technical discussions on The Graph’s .
