r/LocalLLaMA · · 2 min read

Intel releases OpenVINO 2026.4

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

UPDATE https://github.com/ggml-org/llama.cpp/pull/29009 MERGED

More Gen AI coverage and frameworks integrations to minimize code changes

  • New models supported:
    • On CPU: Gemma-3n
    • On CPU, GPU: Kokoro-82M, Qwen3-VL-4B with eagle3, Qwen3-ASR, Muse Glimmer 30B, Qwen 3.8 27B, Gemma4 12B; Hy-MT2-1.8B, DeepSeek OCR-2, Granite 4.0 H Micro
    • On NPUs: FLUX.2-Klein 4B, Kokoro 82M
  • Additional CPU and GPU-enabled models available as early releases: Qwen-image, Z-Image-Turbo, Granite 4.0 H Tiny, Fun-ASR-Nano, LFM2.5-8B-A1B, MiniCPM5-2B, RF-DETR, BGE Reranker-V2-M3, BGE M3

Broader LLM model support and more model compression techniques

  • OpenVINO™ GenAI adds Multi-Token Prediction (MTP) speculative decoding for Gemma 4, Qwen 3.5, and Qwen 3.6, on CPUs & GPUs enabling higher throughput and lower latency without sacrificing accuracy.
  • With Tree Drafting (Top-K) for EAGLE3 now supported in OpenVINO™ GenAI, developers can unlock higher throughput on VLM pipelines compared to Chain Drafting (Top-1).
  • Xe3 integrated graphics optimizations in OpenVINO™ GenAI improve AI inference performance for Gemma 4 models processing long-context inputs on Intel® Core™ Ultra Series 3 processors.
  • OpenVINO™ extends Instrumentation and Tracing Technology (ITT) profiling support to the NPU enabling developers to use Intel® VTune™ Profiler to analyze CPU, GPU, and NPU execution through one consistent toolchain.
  • Preview: OpenVINO™ GenAI introduces DFlash acceleration for Qwen models on GPUs, and visual-token support to reduce latency and speed up GenAI pipelines on Intel® Core™ Ultra Series 3 processors.

More portability and performance to run AI at the edge, in the cloud or locally

  • OpenVINO™ GenAI introduces support for ASRPipeline in Node.js, enabling JavaScript developers to run automatic speech recognition (such as Whisper and Qwen3‑ASR) with streaming and performance metrics using a pipeline API similar to C++ and Python.
  • Preview: Bounded dynamic-shape support on NPUs now extends to vision workloads such as image super-resolution. This capability has been validated with the ESPCN model.
  • Preview: OpenVINO™ Model Server now includes preview support for idle model management, which unloads models when they are not in use to reduce memory usage.
  • OpenVINO™ Model Server adds support for new agentic models such as Muse Glimmer 30B and Qwen3.8 27B.

https://docs.openvino.ai/2026/about-openvino/release-notes-openvino.html

https://docs.openvino.ai/2026/about-openvino/release-notes-openvino/system-requirements.html

submitted by /u/jacek2023
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA