Inference
2 articles

AI
What Are Edge-Native MoE Serving Engines for AI?
A new generation of edge-native serving engines, exemplified by MoEless, slashes AI inference latency by 43% and operational costs by 84% compared to existing solutions.
Arjun Mehta·August 23, 2026

Startups
Infinity AI raises $15M to challenge Nvidia's AI dominance
In a single day, Infinity AI's research agent, Ignition, boosted the inference throughput for the Qwen3-8B model from approximately 1,400 tokens per second to over 20,000 tokens per second.
Diego Navarro·July 20, 2026