#24Models
AI动态:Quantization Aware Healing: a compressed, 4 bit model that outperforms its full precision original
据 huggingface.co 消息,Quantization Aware Healing: a compressed, 4 bit model that outperforms its full precision original。目前来源给出的信息较短,后续还需要继续观察官方公告和产品文档。这条动态值得关注,因为它把模型、工具或基础设施的变化落到了更具体的产品和业务场景。

#26Infrastructure
AI动态:OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
据 techcrunch.com 消息,OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show。Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently.

#27Infrastructure
AI动态:Jalapeño’s first results show industry leading speed and efficiency in AI inference
据 openai.com 消息,Jalapeño’s first results show industry leading speed and efficiency in AI inference。Jalapeño is a custom inference chip from OpenAI that delivers faster, more power efficient AI inference, with higher throughput and lower la.








