LM Studio announced a day-0 partnership with Inco to bring the Splash inference engine into LM Studio. They cite up to about 144 tok/s for Qwen3.8-27B on M5 Max, with setup docs at lmstudio.ai/blog/splash-engine. Splash was already open-sourced by Inco; this puts the Apple-silicon-optimized stack inside LM Studio so local desktop users do not have to assemble the engine separately.

Key Takeaways

  • ✓LM Studio integrates Inco Splash on day 0 for one-click local desktop use.
  • ✓Claims ~144 tok/s for Qwen3.8-27B on M5 Max, with an official setup post.
  • ✓Moves Apple-silicon-optimized inference into a mainstream local client distribution path.
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.