DSpark Benchmark Result on Deepseek v4 Flash 0731
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF
TensorSharp is an native open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. Github repo: https://github.com/zhongkaifu/TensorSharp Thank you for checking out it and starring the project! Any feedback is really appreicated. [link] [comments] |
More from r/LocalLLaMA
-
Hidden Reasoning from Claude and GPT are Decoded, and it is interesting
Aug 12
-
According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1
Aug 12
-
It's the final countdown, baby! Qwen is out in just over 7 hours!
Aug 12
-
FYI: Muse Glimmer Chat Template Got Updated Recently
Aug 12
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.