Majestic Labs, a startup, has announced its Prometheus server concept aimed at solving the memory-bound GPU problem in AI inference. The company proposes replacing costly GPUs with its own Ignite AI Processing Units (AIUs) to eliminate the need for high-bandwidth memory (HBM) in KV caching schemes.
According to the company, the Prometheus server architecture is designed to address the inefficiencies of current AI hardware, where GPUs are limited by memory bandwidth. By using AIUs, Majestic Labs claims to reduce latency and power consumption while maintaining performance for large language models.
As of July 2026, the Prometheus server is still in development, with no confirmed release date. Majestic Labs has not disclosed pricing or performance benchmarks, but early prototypes are being tested with select partners.