French AI startups Kog betting software acceleration GPU inference
TechCrunch
20h ago
Ai Focus
French AI startup Kog improves GPU inference speed through software optimization, plans to advance Series A financing after verifying model acceleration.
Helpful
No.Help

AI 推理提速竞争正在升温。法国初创公司 Kog 没有选择自研芯片,而是押注通过底层软件优化,进一步释放企业现有数据中心 GPU 的推理能力。

先瞄准企业高时延场景

这家公司在今年 5 月因一项技术预览受到关注。Kog 当时展示,使用 AMD MI300X 和 Nvidia H200 等标准数据中心 GPU,也能实现极快的单请求解码速度。公司随后称,这一展示带来了约 200 个明确的商业线索。

Kog 目前优先关注对响应速度要求较高的专业工作流,尤其是软件工程场景。按照公司说法,一些重度代码生成用户等待结果的时间可能长达数小时,推理延迟已成为影响使用体验和成本的重要问题。

公司还在与部分设计合作伙伴推进应用落地。这类客户允许用户通过提示词生成游戏或应用,若推理速度提升,通常意味着更高的转化效率和收入空间。

小模型已跑出 3000 TPS

Kog 此前提出“让大模型推理提速 30 倍”的目标,但目前公开演示仍主要基于小模型。其展示使用的是约 20 亿参数的 Laneformer 2B,单请求推理速度达到每秒 3000 个 token。该模型现已开源。

不过,市场需求并未停留在小模型微调。Kog 表示,潜在客户并不准备把重点放在小模型上,因此公司自发布演示后,已将主要精力转向更大模型的加速开发,以匹配实际需求。

9 月前后验证大模型能力

Kog 认为,GPU 在推理解码上的潜力仍未被完全释放。公司首席执行官 Gaël Delalleau 表示,随着新一代 GPU 内存带宽持续提升,软件层仍有较大优化空间。

这种方法的代价是研发周期较长。Kog 表示,每适配一款新 GPU,团队往往需要投入数周甚至数月做硬件层面的研究。以目前 11 人团队规模来看,公司短期内可支持的芯片数量仍然有限。

Kog 计划未来把这套方法逐步接入基于智能体的研发流程,以支持更多芯片和模型。眼下更关键的节点,是先证明这一路线能在大模型上成立。Delalleau 预计,公司将在 9 月前后完成首个主要模型约 10 倍提速的实现,并据此展示客户进展,推动后续 A 轮融资。

Tip
$0
Like
0
Save
0
Views 67
HKWDB reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Model replaceability, tool pluggability: DeepSeek pushes the competition between Agent from "brain" to "operating system"
On August 14th, DeepSeek open-sourced the Agent framework DeepSeek Harness. Instead of creating a closed product that can only connect to DeepSeek models, they designed components such as models, tools, storage, and sandboxes as plugins, and released them under the MIT license. Developers can integrate models from other companies or replace the execution environment and memory system. For a company that is famous for its foundational models, this choice is quite significant: DeepSeek is shifting its competitive focus from "providing the smartest brains" to "controlling how agents complete tasks."
币界网
·2026-08-15 14:00:28
106
NVIDIA Discloses That Its Holdings of SpaceX Shares Are Worth Approximately $21 Billion
NVIDIA disclosed that its holdings of SpaceX shares were valued at approximately $21 billion at the end of the second quarter, with these holdings coming from investments in xAI; SpaceX also expanded its AI chip collaboration with NVIDIA.
CNBC
·2026-08-15 05:57:16
82
web3: Z.ai Launches GLM-5.3 Programming Model Aiming at the Open-Source Weight Market
Z.ai launches GLM-5.3, focusing on programming capabilities, lower token consumption, and subsequent open-weight releases.
Coinpaper
·2026-08-15 05:17:52
66
web3: Reddit to be included in the S&P 500, stock price surges after the market close
Reddit will be included in the S&P 500 on August 18th, and the news drove the stock price up by about 11% after the market closed. Short-term technical indicators show a strong trend, but there are already signs of overheating.
The Cryptonomist
·2026-08-15 05:06:41
44
California Approves High-Speed Testing of Autonomous Trucks
California issues autonomous truck testing permits to Aurora and Kodiak; the ban on heavy trucks on the road for testing is hereby lifted.
TechCrunch
·2026-08-15 04:44:58
69
View More