Substreams Extended Blocks Data: Data Your RPC Provider Can't Give You

Most applications run on a thin slice of what a blockchain records. A team queries an RPC endpoint, pulls blocks, transactions, receipts, and event logs, and builds on that. It works until the moment the data you need was never in a log. A router that moves funds through an internal call, a contract that changes state without emitting an event, a precise native-balance history: none of that shows up cleanly in a standard RPC response. To get it, developers reach for trace endpoints that are slow, metered, and often switched off.

Substreams take a different route. Firehose extracts every block once and stores it as flat files that Substreams reads in parallel, so the data arrives without the rate limits and repeated polling of live RPC. On EVM chains, that extraction comes in two shapes: the base block model and the extended block model. Which one a chain has decides how much of the chain you can work with. This post explains the difference, and why extended support gives you data that the RPC most developers use today cannot.

Reach out to The Graph Foundation to learn more about Substreams Extended Blocks support

Two Ways Chains Get Substreams Support

A chain can be onboarded to Substreams through two integration paths, and each produces a different block model.

A base block comes from an RPC Poller integration. Substreams polls a standard RPC endpoint and captures what that endpoint exposes: blocks, transactions, receipts, and logs. Because it needs only a normal full node for live data and an archive node for history, a base integration is fast to stand up. A new chain can have Substreams support without shipping custom node software.

An extended block comes from full node instrumentation. Instead of polling an interface, an instrumented node emits the complete internal record of each block as it executes: balance changes, internal calls, storage changes, and the rest of the execution detail. This path takes more work to set up, and it produces the richest data model The Graph can serve for an EVM chain.

The same Protobuf structure carries both. Each block records its own detail level, so a Substreams module can tell whether it is reading base or extended data. The practical effect for a developer is simple. On a base block, the fields that only instrumentation can produce come back empty. As the Substreams documentation puts it, "The data missing in the Base Block makes the corresponding Protobuf field empty. For example, if you try to read internal calls on a Base Block, the list will be empty."

What Lives in an Extended Block

A few categories of data separate an extended block from a base block. Each one answers a question that logs and receipts alone cannot.

Internal calls: When a contract calls another contract, that inner call is not a top-level transaction and does not appear in the transaction list. A DeFi router that splits a trade across pools, a multisend that pays a hundred addresses, a proxy that forwards every call to an implementation: the value moves in these internal calls. An extended block gives you the full call tree for every transaction. A base block gives you the outer transaction and its logs, and nothing beneath.

Balance changes: An extended block records every change to a native-token balance with the reason for each change. You can reconstruct an exact balance history for any address without re-deriving it from transactions and gas math, and without missing the balance effects of internal calls.

Storage changes: Contract state that never emits an event is invisible to log-based indexing. Plenty of contracts update storage without a matching event, either by design or by omission. An extended block captures the storage slots that changed in each block, so you can index and verify state directly rather than inferring it. It also records the Keccak preimages behind those slots, so a hashed storage key can be walked back to its original mapping key and slot when you know the contract's storage layout.

Code changes: When a contract is deployed or removed, the account's code changes. An extended block records these code changes, so you can track every contract deployment directly, including the contracts a factory deploys through internal calls (CREATE and CREATE2) that never appear as a top-level transaction. Read alongside the full call tree, this lets you follow deployed contracts across the whole transaction.

An extended block also records nonce changes for each account. A base block still carries the data most event-based indexing depends on: transactions, receipts, and logs are all there. What it does not carry is the state-level detail above.

Why Extended Blocks Beat Your RPC Provider

The base model already fixes the delivery problems of live RPC. Firehose extracts each block a single time, so Substreams replays history at high speed instead of hammering an endpoint request by request. There are no per-call rate limits, no missed events during reconnects, and the same input always produces the same output. For a team that today stitches together eth_getLogs and eth_getBlockByNumber against a metered provider, even a base-block Substream is a step up in speed and reliability.

Extended blocks add the part RPC providers make hardest to get. Internal calls, balance changes, storage diffs, and contract code changes are not in the standard JSON-RPC surface. To pull them, you have to call trace methods such as debug_traceTransaction or the trace_ namespace, or run state-diff tracers across archive state. Those endpoints are expensive, slow, and heavily rate-limited. Many providers gate them behind premium tiers or disable them entirely, and not every node exposes them. Building a reliable pipeline on top of them is a project in itself.

An extended block has already captured that data at execution time, from an instrumented node, for every block. Substreams then serve it with the same parallel replay, determinism, and reorg awareness as everything else. The data that would cost you a premium trace plan and a fragile scraping layer is simply a field on the block.

What Extended Blocks Support Lets Developers Build

The three data types map directly to things you can build on an extended chain that are impractical on base blocks or raw RPC.

Full call-tree data makes contract accounting correct. If you index a router, an aggregator, or any contract that moves value through internal calls, extended blocks let you attribute every transfer instead of guessing from top-level transactions. The same data feeds MEV analysis, trace-level debugging, and any product that needs to see what a transaction did, not just that it happened.

Balance-change data makes native-token history exact. Wallets, portfolio trackers, accounting tools, and tax reporting can read balances at any block without reconstructing them, and without the gaps that internal transfers leave in a transaction-only view.

Storage-change data makes event-poor contracts indexable. When a protocol changes state without emitting the event you need, extended blocks let you read the state change straight from storage. You can also verify that indexed state matches on-chain state, which matters for compliance-grade and audit-grade data.

Code-change data makes contract deployment traceable. If you track factories or need a live registry of deployed contracts, extended blocks surface every deployment, including the ones created through internal calls that a top-level transaction list never shows.

When Base Blocks are Enough

Extended is not a strict upgrade for every job. A large share of Subgraphs and Substreams only need events, transactions, and receipts, and a base block delivers all of those with the same speed and reliability advantages over RPC. If your indexing logic keys off logs, a base chain serves you well, and the base model is the reason so many chains have Substreams support at all. The base integration path lowers the bar to onboard a network, so coverage grows faster across the many EVM chains launching today.

The decision comes down to the data your use case touches. If you need internal calls, native-balance history, or storage-level state, you need a chain on the extended model. If you index events, base is enough.

How to Check a Chain's Support Model

The Graph’s Supported Networks page, powered by the ecosystem’s networks registry, is the canonical source for which model a chain has. Every EVM network entry carries an evmExtendedModel flag and a blockFeatures field that reads base, extended, or extended@<block>.

As of September 2026, 17 EVM mainnets and 11 EVM testnets carry the extended block model:

NetworkNetwork IDThe Graph Market Support
Arbitrum One Mainneteip155:42161StreamingFast, Pinax
Arbitrum Sepolia Testneteip155:421614Pinax
Arc Mainneteip155:5042Pinax
Base Mainneteip155:8453StreamingFast, Pinax
Base Sepolia Testneteip155:84532Pinax
BNB Smart Chain Mainneteip155:56StreamingFast, Pinax
BNB Smart Chain Chapel Testneteip155:97Pinax
Ethereum Mainneteip155:1StreamingFast, Pinax, DataNexus
Ethereum Hoodi Testneteip155:560048StreamingFast, Pinax
Ethereum Sepolia Testneteip155:11155111StreamingFast, Pinax
HyperEVM Mainneteip155:999Pinax
InjectiveEVM Mainneteip155:1776StreamingFast
InjectiveEVM Testneteip155:1439StreamingFast
Ink Mainneteip155:57073Pinax
Linea Mainneteip155:59144Pinax
Linea Sepolia Testneteip155:59141Pinax
Monad Mainneteip155:143StreamingFast
Optimism (OP Mainnet)eip155:10StreamingFast, Pinax, DataNexus
OP Sepolia Testneteip155:11155420Pinax
Polygon Mainneteip155:137StreamingFast, Pinax
Polygon Amoy Testneteip155:80002Pinax
Robinhood Chain Mainneteip155:4663StreamingFast, Pinax, DataNexus
Soneium Mainneteip155:1868Pinax, DataNexus
Soneium Minato Testneteip155:1946Pinax
Unichain Mainneteip155:130StreamingFast, Pinax, DataNexus
Unichain Sepolia Testneteip155:1301Pinax
World Chain Mainneteip155:480StreamingFast
Zora Mainneteip155:7777777Pinax, DataNexus

Two of these, Arbitrum One and OP Mainnet, are marked hybrid: they provide extended data from a specific block height onward, and the registry lists the exact cutover so you know where the richer fields begin.

The base model covers a larger and faster-growing set of more chains – among them Blast, Celo, Fantom, Gnosis, MegaETH, and Tempo. These counts move as chains are added, so check the registry for the current state.

Non-EVM chains sit outside this base-versus-extended split. Networks like Solana, NEAR, and Bitcoin have their own native block models in Substreams, each shaped to what the chain records.

Core Takeaway

Base blocks give you a faster, more reliable version of the data you already pull from RPC. Extended blocks give you data RPC keeps behind trace endpoints or never exposes at all: the internal calls, balance changes, storage diffs, and code changes that let you account for value correctly, track balances exactly, index state that no event reports, and follow every contract deployment. Before you build on a chain, check its model in the registry, and match it to the data your product needs. If you don’t see your chain, reach out to The Graph Foundation to talk about an integration or request the chain you need.

About The Graph

The Graph is a suite of blockchain data infrastructure products that extract, process, and deliver scalable blockchain data solutions across 60+ networks. The Graph enables application developers, data analysts, AI agents, and enterprise teams that need structured, real-time access to blockchain data. Products include Subgraphs, Firehose, Substreams, and Amp. As of early 2026, The Graph has served over 1.27 trillion queries to more than 75,000 projects, powered by a network of independent Indexers around the world.

Follow The Graph on X, LinkedIn, Instagram, and Reddit. Join the community on The Graph’s Telegram, join technical discussions on The Graph’s Discord.


Categories
Graph UpdatesRecommended
Published
September 28, 2026

The Graph Foundation

View all blog posts⁠