vLLM v0.28.0 has been released, featuring a major update with 584 commits. Key highlights include performance optimizations for Kimi-K3, sparse MLA support for DeepSeek V4, expanded features in Model Runner V2, and extensions to tiered KV cache offloading.
On August 26, 2026, NVIDIA released the FP4 (4-bit) quantized model "nvidia/Qwen3.8-2.4T-A95B-NVFP4" of Qwen3.8-2.4T-A95B on Hugging Face. Quantized via ModelOpt, it is provided as a model for text generation and conversational use.
At Hot Chips, OpenAI announced that its AI inference chip 'Jalapeno,' co-developed with Broadcom, demonstrated performance surpassing Nvidia's GB300 in testing.
OpenAI presented details of Jalapeño, an inference-dedicated chip co-developed with Broadcom, at Hot Chips. Verification by SemiAnalysis indicates that it outperforms competing chips in throughput per watt.
Fixes truncation in floating-point divmod and NaN dropping in median. Adds fused attention paths for NAX devices and memory read optimizations for gqa-8 decoding.
OpenAI has announced the first results for its custom inference chip, Jalapeño, demonstrating high-throughput, low-latency, and energy-efficient inference performance for modern AI models.
On August 25, Z.ai released GLM-5.3, a text generation model utilizing the glm_moe_dsa architecture, on Hugging Face. The model supports English and Chinese and is accompanied by an arXiv paper.
On August 25, 2026, OpenAI announced the suspension of accounts originating from Russia that were using AI to promote fake Israeli think tanks and a pro-Russian "sovereignty" index.
On August 25, 2026, OpenAI announced the 'Admin plugin' for ChatGPT Work and Codex, enabling workspace usage analytics, member and permission management, restriction adjustments, and handling of administrative requests.
OpenAI has announced a price reduction for GPT-5.6 Sol until November 21. Additionally, the company has stopped new access to its fine-tuning platform and established a support period for existing users.
On August 21, 2026, NVIDIA announced that its agent foundation for long-term tasks, "AVO," achieved 100.00 RHAE on the ARC-AGI-3 public set using Claude Opus 5.
Since this summer, Amazon has changed its order confirmation emails to remove specific product names and display only category names. This is seen as a strategic move to adapt to an era where AI agents handle personal data.
Court filings from a federal bankruptcy court reveal that Google will acquire business data, including emails and chats from the bankrupt Spirit Airlines, for $10 million.
Qwen3.8 27B (xhigh) recorded a score of 52 on the Artificial Analysis Intelligence Index. Compared to models of similar size, it is characterized by higher costs and slower inference speeds.
A report on achieving inference for Qwen3.8 27B with a 262,144-token context at an average of 50.44 tok/s on an NVIDIA RTX PRO 4000 Blackwell SFF (24 GB).
Roboflow announced the results of its proprietary benchmark evaluating the visual capabilities of the GPT-5.6 series on July 16. The report indicates that Sol has significantly improved in object detection and counting, demonstrating the strongest VLM performance in OpenAI's history.
Unsloth has released GGUF files for Qwen3.8-27B. The company's Dynamic 3.0 quantization achieves up to a 10% improvement in top-1% accuracy compared to other providers at equivalent sizes.
The Qwen team has released the FP8 quantized build and configuration files for Qwen3.8-27B, an open-weight model with 27B parameters. They also announced that a hosted API featuring a standard 1 million token context will soon be available via Qwen Cloud.
During the solar eclipse on August 12, 2026, the AI image processing of the Xiaomi 17 Ultra malfunctioned, rendering lunar craters and terrain onto the sun. Frandroid confirmed this through comparative photography using four different devices.