AI's Next Challenge: Building a Data Layer for the Web
The rapid expansion of artificial intelligence is creating a pressing need for vast amounts of data. However, much of the information on the web remains locked behind technical barriers or is unstructured, making it difficult for AI models to use effectively. The web was never designed for automated discovery and retrieval, posing a significant hurdle for enterprises seeking to capitalize on AI's potential.
To address this, a new web data infrastructure layer is emerging. This layer must navigate millions of domains and billions of new URLs created weekly, delivering real-time information while overcoming technical obstacles. Or Lenchner, CEO of Bright Data, notes that there is far more data available than most realize, but accessing it requires sophisticated systems capable of handling diverse formats, languages, and access rules.
AI performance now depends not just on model architecture but on the ability to retrieve fresh, relevant, and trustworthy data in real time. Static training data is no longer sufficient for tracking dynamic market trends, competitor pricing, or consumer sentiment. Without real-time context, AI outputs can be stale, leading to poor decisions and eroding user trust. A recent survey found that 56% of AI practitioners believe real-time web data is essential for improving trust in AI outputs.
Despite advances like retrieval-augmented generation, many AI systems still struggle to deliver current and contextually relevant outputs. According to Gartner, 60% of AI projects lacking AI-ready data may be abandoned. As Lenchner emphasizes, retrieving data at scale is not enough; it must also be done in real time to meet user expectations and reduce latency. The next frontier of AI depends on overcoming these infrastructure challenges.