On-chain activity
Inference API
Inference API is a serverless inference tool for running AI models. It helps users process language and vision tasks through a global GPU network.
Inference news, features & analysis
Matched from published articles, podcasts, and talks using the project name, token name, or token symbol.
-
THEA Raises $8M to Build AI Inference Coordination and Settlement Layer on Solana
THEA, a predictive behavioral AI company founded in 2024, has raised $8 million in strategic funding to build a coordination and settlement layer for AI inference on Solana and expand its existing AI infrastructure. ... What the company is building on Solana is a routing, accounting, and settlement layer, one that coordinates inference requests and records transactions as they occur rather than running the inference itself.
-
Arcium Launches Blackthorn: Encrypted AI Inference on Any NVIDIA GPU via MPC
[[PROJECT:236]] has launched Blackthorn, a confidential AI inference framework that enables any NVIDIA H100, H200, or B200 GPU to process AI workloads with inputs, model weights, and outputs remaining fully encrypted throughout computation. ... The framework targets this constraint by ensuring neither the cloud provider, the operator, nor the Arcium network itself can see raw data at any point during inference.
-
Ship or Die at Accelerate 2025: Lightning Talk: Pipe Network
Pipe Network is set to disrupt the content delivery network (CDN) industry with its innovative blockchain-powered solution, promising ultra-low latency and AI edge inference capabilities that could reshape the future of internet content delivery. ... The platform is already attracting attention for its potential applications in AI, particularly for edge inference, and streaming video.
Inference
Inference (formerly Kuzco) is a Solana-based decentralized GPU network that aggregates idle compute from contributors worldwide to deliver affordable, OpenAI-compatible large language model inference to developers and AI teams. The project launched in March 2024 under the name Kuzco, building a distributed GPU cluster on the Solana blockchain for LLM inference. By early summer 2024, the team rebranded to Inference.net to signal a broader ambition: not just a GPU cluster, but a structured inference network with protocol epochs, staking mechanics, and a production-grade developer API. The Solana slug kuzco remains from the original launch. Running large language models in production is expensive. Centralized cloud providers charge premium GPU rates that create barriers for independent developers, startups, and research teams. At the same time, enormous quantities of GPU capacity across data centers, gaming rigs, and consumer workstations sits idle for significant portions of each day. Inference bridges these two realities by building a marketplace on Solana: workers contribute spare GPU time, developers pay for inference at rates the project claims reach up to 90% below centralized alternatives. For GPU contributors, users install Inference worker software on NVIDIA GPUs or Apple M-series chips, then register on-chain via Solana. Once active, the worker receives inference job assignments from the network coordinator. Results are submitted back on-chain, providing a verifiable proof-of-work trail that underpins the reward system. Workers earn $INT points proportional to their throughput and uptime, which convert to $INT tokens at the Token Generation Event (TGE). For developers, the platform exposes an OpenAI-compatible REST API, meaning any application already using the OpenAI SDK can point to Inference with minimal code changes, typically a single base URL swap. The API routes requests to available workers across the global node pool, handling load balancing, worker selection, and failover transparently. The Solana blockchain provides the coordination layer: on-chain job assignment, transparent reward ledgers, and eventual staking mechanics. Inference supports a range of open-source large language models, including Llama 3, Llama 3.3, Mistral, DeepSeek R1, and Phi3. The open-source focus allows the network to serve models that run distributed across heterogeneous hardware, unlike proprietary models that require controlled, monolithic environments. By the time of Devnet Epoch 2 stabilization, the network had connected approximately 5,200 GPU devices and was sustaining around 4,400 inference requests per minute. The network had previously demonstrated peaks of 25,000 requests per minute before the team throttled back to address stability. The team has cited a global estimate of 1.94 billion tokens per second of unused inference capacity, representing approximately $12.2 billion in annualized market value at prevailing rates. Inference has structured its pre-mainnet period into sequential devnet epochs. Epoch 1 established the basic worker and coordinator mechanics. Epoch 2 stabilized the network after scaling experiments revealed instability at higher throughput. Epoch 3, launched in June 2025, introduced changes targeting network scalability, improved operator experience, and tighter economic alignment between worker incentives and network demand. Test tokens used during devnet epochs carry no monetary value and convert to mainnet $INT tokens on a schedule to be announced at TGE. The project native token is $INT, reflected in the tagline CONNECT YOUR GPU EARN $INT POINTS. The staking protocol, currently in Solana Devnet testing, will govern how operators stake $INT to participate as workers, with reward mechanisms tied to uptime and reliability metrics. Inference raised $11.8 million in a Seed round in October 2025, led by Andreessen Horowitz CSX (a16z CSX) and Multicoin Capital. Additional participants included Anatoly Yakovenko (co-founder of Solana), Mechanism Capital, Founders Inc., Topology, Chaotic Capital, and Frictionless Capital. The involvement of Yakovenko and Multicoin signals meaningful alignment between the Inference network and the broader Solana infrastructure thesis. The project is led by CEO Sam Hogan and developed under Context Labs (GitHub: context-labs). Inference competes in the emerging Decentralized Physical Infrastructure Network (DePIN) segment for AI compute, alongside projects like io.net and Akash Network. Its distinguishing features are the tight Solana integration for on-chain coordination, OpenAI API compatibility that reduces developer friction, and a focus on inference workloads specifically rather than training or general compute rental. The combination of a low integration barrier for developers and a straightforward monetization path for GPU owners positions Inference as a two-sided marketplace at the intersection of the DePIN and AI infrastructure trends.
Contents
Solana Token Markets