
CASE meyco3VISION
Okay so Jev can actually do computer use really well
Without any screenshots, or LLMs and no Pixels leave my mac
I dont even read the Dom elements
A local CoreML model segments every button and UI element on screen.
On-device OCR reads the labels. That text is all Jev gets.
It returns a probability across those elements and tells me the best one to click.
Then it clicks, re-runs detection, and decides again. In a loop until the goal is done.
~90ms per decision. Faster than any LLM computer use I've tried.
Blazing fast computer use, without any latency
@typesafeai is building something really interesting译文
行吧,Jev 操作电脑是真的行 不用截图,不用大模型,一个像素都没离开我的 Mac 我连 DOM 元素都不读 本地一个 CoreML 模型把屏幕上每个按钮和界面元素切出来。 端上的 OCR 读出它们的文字。Jev 拿到的就只有这些文字。 它给这些元素返回一组概率,告诉我点哪个最好。 然后点下去、重新识别、再判断一次。循环到目标完成为止。 每个判断约 90 毫秒。比我试过的任何大模型电脑操作都快。 快到没有延迟感的电脑操作 @typesafeai 在做的东西真的有意思
111.6K
浏览
1.6K
点赞
1.4K
收藏
85
转发
引用的帖子
Diogo Almeida @CompleteSkeptic
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x https://t.co/JSybNG2BKJ
34.6M 浏览原帖 ↗
外链与标签
@typesafeai
整理于 Sep 18, 2026, 7:26 PM
更多视觉案例



