Tech Blog

An Evaluation That Trusts No Single Number: Building a Sovereign LLM-Judge Service on Binary Decomposition and Deterministic Gates

Using an LLM as a grader, the practice known as LLM-as-a-judge, is now the default in model development, but the evidence that piled up through 2026 shows that a scalar judge producing a single score is fragile to prompt wording and answer position, drifts toward the middle of the scale, and collapses to coin-flip reliability against adversarial inputs.

Editing Video With a Coding Agent: A Look Inside the video-use Skill

Shared by midudev and quickly making the rounds, browser-use’s video-use is a free, open-source skill: drop raw footage into a folder, type one sentence, and a coding agent handles cutting, filler removal, subtitles, color grading, animation, and rendering.