The Next AI Breakthrough May Depend on a New Web Data Layer
The rapid expansion of artificial intelligence is hitting a critical bottleneck: access to real-time, structured web data. While AI models have advanced significantly, the internet itself was never built for automated data retrieval at scale. Much of the information needed to train and ground AI systems remains blocked, unstructured, or outdated. To overcome this, experts say a new infrastructure layer is emerging—one designed to help AI models navigate hundreds of millions of domains and billions of new URLs created each week.
“The data suggests there's far more data out there,” says Or Lenchner, CEO of Bright Data. “Think of the universe: It's out there, but you don't know what you don't know.” This infrastructure must handle millions of simultaneous interactions across websites varying by geography, language, and access rules. Without it, AI systems risk relying on stale snapshots, leading to poor decisions and eroded user trust.
Speed and freshness are now essential. Real-time data retrieval reduces hallucinations and grounds AI outputs in current, verifiable information. A recent survey found that 56% of AI practitioners believe businesses need access to live web data to improve trust in AI outputs. Yet many systems still struggle, even with retrieval-augmented generation (RAG). According to Gartner, 60% of AI projects lacking AI-ready data will be abandoned by year’s end.
“If it can't retrieve real-time information, it lacks context,” Lenchner says. “In a business setting, that's not acceptable anymore. Stale answers lead to bad decisions and disappointed consumers.” As organizations race to keep pace with dynamic markets, the ability to access fresh, structured data at scale may define the next wave of AI innovation.