Fedβs Watchdog Finds No Misconduct or Illegal Activity in Costly Renovations

Β© Haiyun Jiang/The New York Times

Β© Haiyun Jiang/The New York Times

Β© Heather Ainsworth for The New York Times
While AI has made plenty of inroads on people's phones and computers, it's largely failed in dedicated devices. But over the next year, two major AI companies, Meta and OpenAI, will attempt to change that. And they're apparently betting on a similar path: testing appetite for physical hardware with their cutesy software agents.
"Hardware is hard" is a tech industry mantra. Dedicated AI devices like the Friend and Humane AI Pin have inspired mainly frustration and backlash, which may only worsen as public anger against AI grows. Meta and OpenAI, however, are both pushing ahead. OpenAI is working with famed ex-Apple designer Jony Ive and pla β¦

Β© Alejandro GodíNez/Clasos/ZUMA Press

Β© Karsten Moran for The New York Times
As a Syrian glassmaker created his wares in Damascus, watching his craft reminded this reporter to slow down and enjoy the moment.
Reddit is further limiting who can use the "Old Reddit" experience as part of its efforts to combat scraping and automated traffic. Reddit recently started forcing users to log in to be able to use Old Reddit, but in the "next few months," the company says that you'll have to be logged in and have used Old Reddit within the last six months from when the change rolls out to be able to continue browsing the platform with the classic experience.
Reddit says the six-month window covers "the vast majority of all Old Reddit users." The announcement didn't go into detail, but it's clearly aimed at blocking not just new account creation for bots, β¦

Β© Paul Ratje for The New York Times

Β© Daniel Heuer/Bloomberg News
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.
We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.
Inference engines today all make a performance tradeoff. They are either:
- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)
Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.
Magnitude is built for maximum performance on your hardware and running local agents:
- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels. - Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance. - Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run. - Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.
Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.
Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:
Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s β 57 tok/s) - 9% faster prefill (466 tok/s β 507 tok/s) - 28% less per-agent memory usage
CUDA (DGX Spark) - 19% faster decode (49 tok/s β 58 tok/s) - 23% faster prefill (2,033 tok/s β 2,507 tok/s) - 27% less per-agent memory usage
Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o
We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:
- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU. - Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better. - Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.
We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!
Comments URL: https://news.ycombinator.com/item?id=49911995
Points: 32
# Comments: 10

Β© Al Drago/Bloomberg News

Β© Janice Chung for WSJ
Amazon deliveries could soon come with a new catch: the delivery driver's smart glasses will be snapping photos of anything they see while they're in use, including people and private property. Bloomberg reports that the glasses could take "several thousand captures in a single driver's typical shift," which Amazon plans to upload to its AI Wellspring platform. When asked if customers will be able to opt-out of this, Amazon's Viraj Chatterjee told Bloomberg, "We haven't thought about that."
Chatterjee claimed Amazon's only using the data from the glasses to "enhance the delivery process," adding, "The intent has never been stuff like surve β¦
The Cornell lawsuit filed by a woman alleging a brutal rape raises questions on how police handle such cases. The majority of police lack adequate training on trauma-informed approaches, experts say.
(Image credit: John Munson/AP)