English

NewsLaya

Laya Ported to Core ML for High-Speed, Energy-Efficient Offline Typed Decisions on Apple Silicon

A developer has released laya-coreml, an open-weight port of Laya that enables typed decisions to run locally on Apple Silicon using Core ML and the Neural Engine (ANE). The implementation allows for offline inference without the need for PyTorch, Transformers, or MLX.

In benchmarks conducted on an M3 Max, short multilingual decision tasks achieved a P50 latency of 4.98 ms and a P95 of 5.31 ms using ANE FP16. The implementation showed significant energy efficiency improvements, measuring 2.78× better whole-system energy per decision compared to compiled MLX FP16. A W8-quantized variant further improved energy efficiency by 3.19×.

The project demonstrates a "Snake" game loop that sustained between 49.1 and 50.0 decisions per second across three 600-step episodes. Unlike autoregressive models, this approach provides probabilities for choices, ordinal scores, and boolean answers without generating tokens or requiring JSON parsing.

The implementation includes tools to convert existing Laya models into Core ML formats and provides an inference wheel for Python. While the ANE-optimized version shows high performance for short queries, the developer notes that the current adapter's performance for longer contexts or continuous full-game speedups relative to MLX requires further validation.

Sources

  1. Laya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second) (Hacker News Frontpage, 2026-09-20)
  2. GitHubリポジトリ