October 23, 2025By Pete Sonsini
Tensormesh: Cutting the Cost of AI Inference by 10x
Our lead investment in the team commercializing LMCache for enterprise inference.
We're thrilled to share our lead investment in Tensormesh, which recently emerged from stealth with a $4.5m seed round. The team's tackling one of the most urgent problems in enterprise AI today: the increasing cost and complexity of running LLMs in production.
The founding team (Junchen Jiang, Yihua Cheng, and Kuntai Du) are pioneers in AI infrastructure. They spent years at the forefront of the field at the University of Chicago, Carnegie Mellon, and UC Berkeley. Their advisors include Ion Stoica, Mike Franklin, and Hui Zhang – and even before starting Tensormesh, their widely-adopted open-source LMCache project was already trusted and integrated by technology heavyweights including Bloomberg, Red Hat, Redis, Tencent, and WEKA.
The timing couldn't be better. The industry demands the kind of high-level optimization that can only come from deep research. LMCache, the team's open-source offering, provided an early map and proved the core concept. Now, Tensormesh enables superior performance with a critical addition: control and enterprise readiness.

When AI gets too expensive to run
Every company scaling AI right now is hitting the same wall: inference costs are destroying budgets.
Compute is expensive; controlling latency is a constant battle. Organizations must either rely on costly third-party APIs or hire huge engineering teams to build complex, custom inference systems in-house. Plus, data sovereignty and compliance demands are only increasing, making reliance on external hosting riskier for companies handling sensitive data.
How Tensormesh stops wasting the model's "thinking"
Tensormesh doesn't just manage inference better; they solve a fundamental inefficiency in how today's models run.
When an LLM performs inference, it generates crucial intermediate computational states we can think of as the model's "thinking." Today's standard systems discard this data due to GPU memory limitations, which means models constantly have to reprocess the same information. That burns compute cycles and slows down the entire application.
Tensormesh instead captures and reuses these intermediate states. Their upgraded LMCache is a distributed caching layer that uses CPU, disk, and GPU resources. When combined with the Tensormesh Production Stack for managing distributed inference, enterprises can find up to 10x faster performance, all while keeping secure control over their models and data.
It's an elegantly straightforward idea with a huge payoff: eliminate redundant work, maximize your existing hardware, and break away from external vendor dependency.
The right technology at the right time
As enterprises deploy conversational interfaces and sophisticated agentic systems with ever-growing context windows, the cost of repeated computation is becoming a massive line item in their AI infrastructure budget.
Tensormesh gives organizations a practical way to deploy cutting-edge AI on infrastructure they already own. The early signal is undeniable: the foundational LMCache has been integrated into Google Kubernetes Engine and incorporated by Nvidia into their own infrastructure tools. This incredible open-source traction is powerful validation of the team's technical vision and its broad industry relevance, proving the core technology before the company even launched.
This is exactly the kind of team and technology we look to support: founders with deep, academically rigorous research, a profound understanding of the problem space, and a solution that creates immediate, clear value for enterprises building real AI systems.
We believe Tensormesh is perfectly positioned to become a foundational component of the entire AI infrastructure stack as the demand for cost-efficient, high-performance, and secure inference grows.
Congratulations to Junchen, Yihua, Kuntai, and the entire Tensormesh team on this major milestone. We are proud to support your journey from the research lab and open-source community to production environments across the industry.
If you're building the next generation of AI infrastructure and wrestling with these challenges, we'd love to hear from you: hello@laude.vc