EmbeddingGemma 2

EmbeddingGemma 2 puts text, code, images, video and audio in one shared 768-dimension vector space, under Apache 2.0, in about 191 MB on a phone. What changed since v1, what the MTEB numbers actually say, and the float16 warning that will quietly ruin your retrieval.

Strata: Run a 125B Model on a 12 GB Gaming GPU

Strata runs Qwen3.8-Flash-Next, a 125B MoE model, on a 12 GB gaming GPU at 94 tokens per second. How the three-tier memory split works, what the self-reported numbers do and do not show, and the licence detail that decides whether you can ship it.

Omarchy: The Linux Desktop Built for AI Agents

Omarchy 4.0.4 wires thirteen coding agents into Arch and Hyprland. Agent spend in the status bar, crashes routed to your agent, 367 self-documenting commands, and a shipped skill file that teaches any agent how the system is built.

GLM-5.3-Flash on Two DGX Sparks: The TensorFold Recipe

Two DGX Sparks, a cable between them, and GLM-5.3-Flash answering on an OpenAI-compatible endpoint with its full million-token window open. The repository publishes enough numbers to decide before you buy the hardware, and one licence detail that decides it for commercial users.

5 Things Your AI Agent Needs From a Generation API

Plenty of services can generate an image. Far fewer can be used by an agent. Your agent can call the endpoint, but can it tell what the call will cost, whether the model is working, and can it pass the result into the next step without glue? Here are the five checks that decide it.

n8n MCP Features Explained: Five Things, One Confusing Name

n8n now has five separate features with MCP in the name, and two of the node names differ by one word. Here is the map: what each one does, which fits your problem, the one registry server self-hosted instances never see, and the validation trap that passes a server which does not exist.

Ruflo

Ruflo runs hierarchical and mesh agent swarms with Raft, Byzantine and Gossip consensus. MIT, 73,553 stars. It calls itself the original agent harness; OpenCode predates it.

claude-mem

claude-mem hooks the session lifecycle, compresses what your agent did and re-injects it next time. Works across many agents. Defaults to hosted memory, so set the local flag.

CodeGraph

CodeGraph parses your repository into a local SQLite knowledge graph and serves it over MCP to nine agents. 88% fewer tool calls, 62% fewer tokens, 100% local.

anydoc

Firecrawl's Rust converter handles 14 formats including legacy Office binaries, at a 4.4 ms median. No local OCR, which is the catch worth knowing.

Jeff

Jeff fine-tunes Qwen3.5 and Gemma 4 into zero-shot decision models scoring 82.0% against Jev's 83.0%, at about 22 ms per decision on local hardware.

Hermes Agent

Nous Research Hermes Agent writes its own skills, refines them in use and models you across sessions. MIT, model-agnostic, seven sandbox backends. And it corrects a claim we made.

Herdr

Herdr owns the terminals your coding agents run in rather than wrapping them. Agent-aware pane state, survives a dropped SSH session, and a socket API agents drive themselves.