On-chain activity
DataHive Staking
DataHive Staking is a Solana validator service where users delegate SOL to earn standard staking rewards plus $DATA airdrops and platform benefits such as task access and point multipliers.
DataHive AI
DataHive AI is a decentralized physical infrastructure network (DePIN) built on Solana that turns ordinary user devices into a globally distributed data collection and labeling pipeline for the AI industry. Founded in October 2024, the project bridges a structural gap in AI development: the need for large volumes of high-quality, rights-cleared training data that centralized vendors like Scale AI and Appen supply at significant cost. DataHive routes that work to its contributor network instead, claiming a 10–20x cost reduction for enterprise buyers.
How It Works
The platform operates through two lightweight clients: a Chrome browser extension and an Android mobile app. Both run passively in the background during normal device use, collecting publicly accessible web content—text, images, video, and dynamic page data—without touching personal information, login credentials, or private browsing history. Because collection routes through real residential devices with unique IP addresses and natural browsing behavior, the network captures JavaScript-rendered pages and live e-commerce data that centralized crawlers routinely miss.
Running both clients simultaneously on the same account provides a combined earning boost of up to 2× the base point generation rate.
Dataset Catalog
DataHive publishes a growing catalog of rights-owned datasets built from its contributor network. Current offerings span several categories:
- Multilingual speech: Arabic multidialect emotional speech with emotion labels and speaker demographics; European languages spoken audio for speech recognition; professional linguist-validated transcriptions.
- E-commerce data: product listings covering titles, brands, pricing, and availability across 1.5 million products; more than 100 million customer ratings and reviews.
- Video and image: a global video dataset of over 1,000 hours with sentiment annotations; an image and photo collection with metadata and contextual tags.
- Entertainment reviews: film reviews spanning 2000–2024 with associated metadata.
All datasets carry full IP ownership and dual-rater quality assurance, positioning them as ready-to-use for AI training workflows.
Rewards and Token Design
Contributor earnings accumulate as points across three separate tracks, all converting to $DATA tokens at the Token Generation Event (TGE):
Data Points are earned through active collection tasks processed while the extension or app is running, measuring the quality and quantity of data contributed.
Hive Points are earned through device uptime—cumulative hours a device remains connected to the network—rewarding sustained participation independent of whether specific tasks are being processed.
SOL Staking Points are earned by delegating SOL to DataHive's Solana validator. This track runs concurrently with standard Solana staking yield, meaning stakers earn both the network's native APY in SOL and additional points toward the $DATA airdrop from a 1 SOL minimum.
The $DATA token is confirmed on Solana. Over 642 million points had been distributed to active participants ahead of TGE, which had not been formally dated as of the project's latest communications.
Funding and Backers
In August 2025, DataHive AI closed a $3.5 million seed round led by 6th Man Ventures (6MV), with participation from Solana Ventures, Alliance DAO, Race Capital, Nural Capital, and Side Door Ventures. Angel investors include Solana co-founders Anatoly Yakovenko and Raj Gokal, and investor Santiago Roel Santos. The backing from Solana's founding team places DataHive inside the ecosystem's core network and signals alignment with Solana's DePIN thesis.
Market Position
DataHive competes in a sector where the cost and ethics of AI training data are under increasing scrutiny. Its core claim is that routing collection through consenting users' devices—rather than centralized scraping operations—produces data that is both higher quality (residential IPs access live, JS-rendered content) and more defensible under emerging AI data regulation.
The project reports paying enterprise customers and published, purchasable datasets before its token launch, distinguishing it from DePIN projects that defer revenue entirely to a post-TGE phase.
Solana Integration
DataHive's choice of Solana as its settlement and reward layer reflects the network's throughput and cost profile. Running a dedicated Solana validator ties the project's operational infrastructure directly to the chain, while SOL staking points create an onramp for existing SOL holders who want exposure to the project's $DATA airdrop without running the extension or app.
Contents
- How It Works
- Dataset Catalog
- Rewards and Token Design
- Funding and Backers
- Market Position
- Solana Integration
Solana Token Markets