AI News List

List of AI News about Modal

Time Details
2026-07-15
21:06
Modal DFlash speculator boosts inference 67%

According to soumithchintala, Modal’s DFlash speculator for Inkling delivers 67% higher throughput and interactivity than MTP on SGLang endpoints.

Source
2026-06-24
11:50
DFlash Boosts Qwen inference 4x with zero loss

According to @_avichawla, DFlash speculative decoding lifted a 122B Qwen model from 250 to 1000+ tokens sec with zero quality loss by parallel drafting.

Source