Unlocked by Data Foundation

This research report has been funded by Data Foundation. By providing this disclosure, we aim to ensure that the research reported in this document is conducted with objectivity and transparency. Blockworks Research makes the following disclosures: 1) Research Funding: The research reported in this document has been funded by Data Foundation. The sponsor may have input on the content of the report, but Blockworks Research maintains editorial control over the final report to retain data accuracy and objectivity. All published reports by Blockworks Research are reviewed by internal independent parties to prevent bias. 2) Researchers submit financial conflict of interest (FCOI) disclosures on a monthly basis that are reviewed by appropriate internal parties. Readers are advised to conduct their own independent research and seek advice of qualified financial advisor before making investment decisions.

The Onchain AI Data Stack

Nick Carpinito

Key Takeaways

  • AI labs have hit the data wall, and the binding constraint has shifted from compute to data labs can attribute and audit. The open internet has been scraped, the remaining high-value data is proprietary or legally encumbered.
  • A vertically segmented onchain data stack has formed across four layers: collection, processing, labeling, and provenance. Contributor networks supply the raw material, processing and labeling networks make it model-ready, and provenance rails make it usable and attributable.
  • Regulation is converting provenance from a feature into a procurement requirement. The EU AI Act's training-data transparency mandate and mounting copyright litigation exposure mean labs increasingly need documented provenance, consent, and attribution rather than cheaper unverified data.
  • Value concentrates at the layers with verified supply, quality enforcement, and provenance. Raw bandwidth and undifferentiated scraping are race-to-the-bottom inputs with thin margins. The defensible positions are the two-sided marketplaces and provenance registries that accumulate a data and reputation moat as they scale.

Subscribe to 0xResearch Newsletter

Introduction

The most expensive input to a frontier AI model has become data whose origin labs can prove and attribute. Frontier labs have effectively exhausted the freely scrapeable internet. What remains is either proprietary and bespoke, or undocumented and legally risky, and the cost of getting provenance wrong runs into the billions in litigation exposure.ai_training_dataset_market_bwr.pngAs of late 2025, more than 50 copyright lawsuits between IP owners and AI developers were pending in US federal courts. The global data monetization market is projected to exceed $7B in 2026 and grow at a 24.1% CAGR to $52B by 2035, but that figure understates the strategic stakes. A bottoms-up read on today's market lands in a similar range. One mid-2026 market map estimates the offchain field for training data and RL environments at ~$8.5B in gross run-rate revenue and ~$96B in private valuations across about 28 companies, led by Scale AI near $1.5B in revenue, Surge AI around $3B, and Mercor above $2B. Labs already pay billions for training data and labeling, and they route that spend to offchain vendors, so the onchain question is not whether demand exists but whether provenance rails capture any of it. Foundation-model labs are now posing the same three questions to data suppliers:

  1. Can you source the data we need at scale?
  2. Can you prove the data’s origin?
  3. Can you ensure the quality of the data?

Crypto is structurally suited to address contributor incentivization and data provenance. Token incentives bootstrap a global contributor base faster and more cheaply than enterprise procurement, and an onchain registry produces a tamper-evident, queryable audit trail by default.

The Data Stack

The onchain data supply chain separates cleanly into four layers, each with a distinct margin profile and competitive structure.

Screenshot 2026-07-20 at 8.52.22 AM.png

The thesis embedded in this taxonomy is that value migrates upward. Layer 1 commodity inputs compete on price and contributor count; Layer 4 provenance infrastructure compounds a data and compliance advantage that is expensive to dislodge.

Collection

The collection layer is the most crowded and the most internally varied. It splits into three sub-models with very different economics.

Bandwidth and Scraping

The largest cohort by user count monetizes idle residential bandwidth, routing it so AI clients can scrape the public web from geographically diverse, residential IP addresses that are harder to rate-limit or block. Projects of this archetype follow a supply-side playbook where users install software, passively route bandwidth, and earn rewards.

Grass is the category leader. It positions itself as a sovereign data rollup, uses zero-knowledge proofs to checkpoint session data onchain for provenance, and reports collecting 1.1K TB of data daily from 8.5M active users as of November 2025.

Nodepay routes idle residential bandwidth into real-time data retrieval for AI, positioning its feed for retrieval-augmented generation rather than the bulk web-scraping corpora Grass targets. It reports ~1.5M active users across 180 countries.

Gradient Network raised ~$10M led by Multicoin with Pantera and Sequoia participation, and layers AI infrastructure (Parallax, Echo, Logits) on top of node participation.

The structural read on this sub-layer is that bandwidth is a commodity. These networks compete on contributor incentives and airdrop mechanics more than on a differentiated product. The conversion of supply-side enthusiasm to demand-side revenue is the test the entire collection layer has yet to pass at scale.

Scraping Subnets

Within the Bittensor ecosystem, two subnets target the same data-sourcing function through a different incentive structure. Miners compete for token emissions by supplying the freshest, most in-demand data rather than the most bandwidth.

  • Data Universe (SN13) is a decentralized data layer that incentivizes miners to scrape real-time social data, primarily X and Reddit, with the team reporting the system has collected ~55B rows to date. Its Gravity product lets buyers specify and prioritize the terms they want scraped, a self-organizing data marketplace.
  • Masa (SN42) focuses on real-time, structured, annotated training data as an open-access layer of the AI stack, positioning explicitly against the data monopolies of incumbent labs.

These subnets matter to the landscape because they demonstrate the emission-driven model for bootstrapping data supply, and because the wider Bittensor ecosystem also hosts the labeling function discussed in Layer 3 (Dojo).

First-Party Contributor Apps

The economically distinct sub-model collects consented, first-party data directly from users, which is harder to replicate than scraped web data and carries clean consent at the point of capture. This is the segment with the strongest moat potential because the data cannot simply be re-scraped by a competitor.

Numo is a consumer contributor app collecting consented voice in underrepresented languages (Bengali, Hindi, Tamil, Telugu, Vietnamese) and registering each contribution onchain on the Data Network at the point of capture. The platform reported 15K+ contributors and 210K+ recordings within weeks of its April 2026 early-access launch. Voice is the initial modality, but the underlying architecture is designed to expand across video, image, movement, and sensor data as demand warrants.

Oto is an independent talk-to-earn app targeting the data type speech-to-speech voice models require, full-duplex conversations with each speaker recorded on a separate channel at high sample rate, which the mixed audio in podcasts or video cannot reconstruct. Users are matched into short, live conversations and paid to talk. By Oto's account, the app has drawn 20K+ users and several thousand hours of audio and has released OtoSpeech, a full-duplex conversational corpus, on Hugging Face.

Miso extends the contributor model to an existing consumer platform rather than a purpose-built data app with its Korean on-demand home-services marketplace. Through the DATA integration, it routes that task workflow toward consented Korean voice data contribution, registering each record onchain with a receipt.

Kled is an opt-in human-data marketplace where users sell personal images, video, and files routed to AI labs. It reports millions of daily uploads against 1.5B+ total uploads and 500K+ total users. Kled now runs on the Data Network as the ecosystem's flagship app, with contributor receipts and verifiable audits. Payouts settle in USDC on DATA Network alongside stablecoin and fiat options, with the payment trail auditable on the same registry.

txs.png

Fluffle is a voice-calling app that pays users to opt in to recording their conversations, generating consented, anonymized voice data that scraped or synthetic datasets cannot replicate. It is a direct analog to Numo on the supply side.

Processing and Quality

Raw contributions arrive unstructured, inconsistently formatted, uneven in quality, and contaminated with scraped, pirated, or synthetic content that labs cannot afford to ingest, because once bad data enters a training pipeline, it is effectively impossible to remove. The processing layer cleans, normalizes, deduplicates, validates authenticity and license, and scores each record for quality before it reaches a buyer.

poseidon_audio_quality_bwr.png

Poseidon is the most directly relevant onchain processing layer, a Data Foundation incubated network that turns raw human contributions into datasets labs can train on, with authenticity screening run up front to filter scraped, synthetic, or altered content. It deploys domain-specific subnetworks (audio, robotics footage, autonomous-vehicle edge cases, multimodal) on Data Network infrastructure on the premise that medical data and robotics data have fundamentally incompatible processing requirements and cannot share one pipeline. Contributors upload raw data, which moves through automated deduplication, AI-assisted annotation, and consensus-based quality validation before being registered on the Data Network and listed on an open marketplace where AI developers access licenses. Poseidon claims dataset coverage for low-resource languages that is often an order of magnitude larger than public alternatives like CommonVoice, FLEURS, and MLS.

newplot.png

Ocean Protocol occupies the adjacent privacy-preserving position. Its V4 architecture separates a base ownership claim (a data NFT) from transferable sub-licenses (Datatokens), and its compute-to-data model lets buyers run training jobs against a dataset without ever downloading the raw files, a design built for sensitive data where the owner will not relinquish a copy.

The defensibility of processing quality is a genuine differentiator, but it is a service layer, and its moat depends on proprietary validation pipelines rather than network effects.

Labeling and Human Feedback

Labeling is the most labor-intensive and, historically, the most centralized layer of the AI data supply chain, dominated offchain by vendors built on low-paid gig labor. It is also where reinforcement learning from human feedback (RLHF), the process of using human judgments to fine-tune model behavior, creates a structural bottleneck where models need large volumes of high-quality human judgment, and quality is hard to enforce across a distributed crowd.

Sapien is the clearest onchain entrant, a decentralized data foundry on Base that connects enterprises with a global contributor base and enforces accuracy through a Proof-of-Quality system combining peer validation and token staking. It reports 1.2M+ registered users and 100M+ tasks completed across 27+ enterprise customers as of mid-2025.

Within Bittensor, Dojo (Tensorplex) targets the same crowdsourced labeling and human-feedback function through subnet emissions.

The offchain incumbents set the benchmark these entrants compete against. Scale AI, Surge AI, and Mercor each run into the billions in annual revenue, so Sapien and Dojo are contesting a market with proven demand rather than building one from scratch.

The labeling layer carries a stronger moat than collection because quality enforcement and contributor reputation compound over time, a labeler base with a proven accuracy record is not easily replicated, and staking-based quality systems raise the cost of low-effort participation. This is the come for the tool, stay for the network dynamic applied to data work. Contributors arrive to earn, and the platform retains them through reputation that is portable only within the network.

Provenance, Attribution, and Audit

The top of the stack answers the question that determines whether any of the data below it can be used: where it came from, who consented to it, and whether that origin can be proven. This is where commodity inputs become institutional-grade assets, and it holds the most durable moats:

  • Data: compounds with every record registered
  • Regulatory: created by compliance mandates
  • Switching costs: grow as a buyer's audit history accumulates on one registry

Two architectures compete to own this layer, approaching it from opposite ends. The Data Foundation builds a top-down provenance and attribution registry. Vana builds a bottom-up market for user-owned data. Both answer the right question, but for different data and different buyers.

The Data Foundation

The Data Foundation (formerly Story Protocol) addresses a structural problem common to every data vertical: a dataset's origin travels separately from the data itself, and that gap widens each time the data changes hands. The underlying network, built by PIP Labs and backed by ~$140M in funding, launched as a purpose-built L1 in February 2025. The Data Foundation focuses on provenance: registering where data came from and who consented to its use, then proving both to buyers who need to trust the source.

newplot_relabeled.png

Trace is the Data Foundation's provenance and audit surface. A contributing app collects data with explicit consent and posts a SHA-256-anchored receipt to Trace over a signed webhook, which Trace batches into a Merkle root and registers via the Data Foundation, appending lifecycle events over time. Each receipt carries a content hash, a consent acceptance, KYC status, media type, and file size, plus a per-record metadata-update history, queryable by labs and regulators while contributor identities stay anonymized and the underlying data stays with the contributing app. The Foundation operates Trace as a whitelabel layer, giving each app a branded trust portal at its own subdomain while aggregating records into a single public registry. This lets labs verify provenance across sources without trusting each one individually. Registered volume is overwhelmingly Kled, with Numo and Oto also live as contributor apps. On June 26, 2026, Numo went live inside Toss (Korea's finance super app, ~30M users), embedding Korean-language voice, image, and video collection directly in the Toss financial app, with payment routed through Toss to contributors. Trace reports 122.9M indexed receipts from 186K contributors and 190.9 TB audited today. The Data Foundation attributes the climb to new daily contributions plus Kled registering a backlog of roughly 1B files accumulated over the prior year.

Confidential Data Rails (CDR) closes the delivery gap. Announced in November 2025 and live on the Aeneid testnet since April 2026, with mainnet targeted for Q3 2026, CDR moves encrypted, high-value data between untrusting parties without exposing it to any intermediary, including the protocol itself. Each Trusted Execution Environments (TEEs) enabled validator holds a fragment of a threshold decryption key. An owner encrypts a file client-side, attaches it to an asset, and defines programmable access conditions, time-windowed access, TEE-restricted compute, or multi-party approval, and the network reconstructs the key only when those conditions are satisfied. CDR is distinct from zero-knowledge proofs. While ZK technology proves data exists without revealing it, CDR controls who can access the data and under what runtime conditions. The two approaches are complementary rather than competitive. Its most consequential application is a training data marketplace where developers access encrypted datasets under enforced conditions, with delivery and access control handled at the protocol layer, satisfying the EU AI Act's Article 53 training-data transparency requirements without exposing the raw data to the buyer.

Vana

Vana approaches the same layer from the ownership side. It is an EVM-compatible, proof-of-stake L1 that went to mainnet in December 2024 with VANA as its native token, built so individuals can pool personal data and retain a persistent proportional claim on the models trained on it rather than taking a one-time payment. The coordination primitive is the Data Liquidity Pool (DLP), a smart contract that instantiates a Data DAO around a specific data source and maps non-fungible personal data to a fungible, dataset-specific token. Contributors submit data, a per-DAO Proof-of-Contribution function scores it for quality (each DAO sets its own parameters, since financial, genetic, and social data carry different quality measures), and contributors receive that DLP's token, which confers governance and profit-share rights. Vana's VRC-20 standard formalizes these data-backed tokens with fixed supply, governance, and liquidity rules, and DLP creation is permissionless because it leverages existing data-export rights that let users extract their data from the platforms that hold it.

The model has produced meaningful breadth: Vana reports 300+ Data DAOs and on the order of 1.3M users, with its developer Playground launching in September 2025, carrying ~12.7M data points from ~1M users. The Reddit Data DAO, aggregating post and comment history from 140k+ users behind the RDAT token, is the most-cited live example.

The contrast frames the layer cleanly. The Data Foundation is the enterprise provenance registry, a documented origin trail plus confidential delivery, aimed at labs that need to source high-value data and prove where it came from. Vana is the consumer data-financialization network, a per-community ownership and pricing mechanism for pooled personal data. They are not direct substitutes, but rather sit at the two ends of the provenance layer, and a mature market plausibly needs both an institutional registry for documented provenance and a sovereignty model for individual data.

Value Accrual

Reading the stack through a tech-investment lens clarifies where the durable businesses sit.

The collection layer's bandwidth cohort exhibits weak, non-exclusive network effects and minimal switching costs. A contributor running Grass can run Nodepay and Gradient simultaneously, and frequently does, so supply is non-rival and undifferentiated. These are the bare-metal equivalents of the data stack, essential but commoditized, competing on incentive design rather than product.

onchain_data_stack_value_ladder_v5.png

First-party contributor apps and labeling networks sit higher because their inputs are harder to replicate and their quality systems compound. A consented voice dataset in an underrepresented language, or a labeler base with a verified accuracy record, cannot be re-scraped or instantly cloned.

Provenance infrastructure is the strongest position for three reasons:

  1. Regulatory Moat: As compliance mandates make verifiable provenance a procurement requirement, the registry that labs already trust becomes the default.
  2. Switching Costs: A buyer's accumulated audit history is sticky, and moving registries means re-establishing provenance from scratch.
  3. Natural Aggregation Layer: The connective tissue that unifies fragmented collection, processing, and labeling outputs into a single verifiable supply.

The fragmentation across dozens of small, incompatible networks is the opportunity. The market breaks down precisely because most players solve one layer, and labs cannot assemble a compliant end-to-end supply from a patchwork of single-layer vendors. Teams that credibly connect sourcing, quality, and provenance will capture the value the individual layers cannot.

Risks

Supply-to-demand conversion: The defining risk across the collection layer is that token-incentivized contributor growth has not been shown to convert into recurring, demand-side revenue from labs at scale across the cohort. Supply is also non-rival, a single contributor runs Grass, Nodepay, and Gradient simultaneously, so node counts overstate differentiated capacity until matched by paying enterprise demand. Recent onchain activity illustrates the gap. The Data Network posted fees near $4,000 per day in mid-July, up sharply from weeks earlier, but that spend comes mostly from contributor apps registering data, not from labs licensing it.rev.pngPoint of Capture: An onchain audit trail records what it is given. If consent is defective at collection, the registry faithfully propagates a defective claim. The legal enforceability of onchain licenses against unwilling third parties also remains largely untested in court; the onchain record does not replace the offchain legal instrument, and a dispute will ultimately resolve under national IP law.

Regulatory Ambiguity: Compliance mandates create the demand thesis, but regulatory treatment of contributor rewards, data-as-a-security questions, and cross-border data transfer rules could raise costs or restrict distribution.

Data Quality: As synthetic content proliferates, distinguishing authentic human data from model-generated or altered data becomes harder and more adversarial. A processing layer that fails to screen contamination passes a latent liability downstream that cannot be reversed once trained on.

Concentration and Counterparty Risk: Several layers depend on a few large contributors or a single marquee marketplace for the bulk of their volume. Concentrated supply accelerates reported growth but raises fragility if a key counterparty's figures, methods, or compliance posture prove weaker than claimed.

Conclusion

The onchain data sector is consolidating around provenance as the binding constraint. The infrastructure to source, process, and label data has reached production readiness across all four layers, but readiness is not adoption. Demand-side conversion has not been shown as recurring lab revenue across the cohort, separating the sector from institutional scale. This is the same gap that has repeatedly separated infrastructure readiness from institutional adoption elsewhere in crypto, and it is the variable that decides whether the provenance thesis pays out.

The competitive structure favors the top of the stack. Commodity bandwidth and undifferentiated scraping will compete on price and contributor incentives indefinitely, while value compounds at the layers that own verified supply, enforce quality, and prove provenance. The fragmentation of the current landscape, dozens of single-layer networks that labs cannot easily stitch into a compliant pipeline, is the structural opening for connective provenance infrastructure to become the sector's aggregation layer.

Three catalysts are most likely to drive the next 12 to 18 months:

  1. The EU AI Act's training-data transparency requirements turn provenance from an advantage into a procurement requirement, the first hard external deadline for adoption.
  2. The migration of large first-party data marketplaces onto verifiable onchain rails would, if the volume claims hold up to scrutiny, move the sector to enterprise-scale supply.
  3. The maturation of confidential data delivery would unlock the high-value private datasets (medical, robotics, proprietary enterprise data) that the current scraped-web supply cannot reach.

The primary risk to all three is that supply-side growth continues to outrun verifiable demand-side revenue. Until the sector can show labs paying recurring fees for provenance-verified data at scale, the data wall remains a thesis about where value should accrue rather than proof of where it has.

This research report has been funded by the Data Foundation. By providing this disclosure, we aim to ensure that the research reported in this document is conducted with objectivity and transparency. Blockworks Research makes the following disclosures: 1) Research Funding: The research reported in this document has been funded by the Data Foundation. The sponsor may have input on the content of the report, but Blockworks Research maintains editorial control over the final report to retain data accuracy and objectivity. All published reports by Blockworks Research are reviewed by internal independent parties to prevent bias. 2) Researchers submit financial conflict of interest (FCOI) disclosures on a monthly basis that are reviewed by appropriate internal parties. Readers are advised to conduct their own independent research and seek the advice of a qualified financial adviser before making any investment decisions.