On October 10, 2026, GitHub's daily trending list included alibaba/open-code-review (Apache-2.0, over 45,000 stars), the same day as the still-climbing morluto/rea this digest covered on October 7. Its latest release, v1.12.13, shipped on October 8 with scan and timeout fixes.
Where it came from
The README says the tool "originated as Alibaba Group's internal official AI code review assistant — over the past two years, it has served tens of thousands of developers and identified millions of code defects." It "reads Git diffs, sends changed files to a configurable LLM via an agent with tool-use capabilities, and generates structured review comments with line-level precision." An ocr scan mode reviews whole files when there is no useful diff. It works with OpenAI- and Anthropic-compatible endpoints and has plugins for Claude Code, Codex, Cursor and Kimi Code.
The design argument
The project's case against simply pointing a general coding agent at a pull request is specific. With "general-purpose agents like Claude Code with Skills for code review," it lists three failures:
- "Incomplete coverage" -- "On larger changesets, agents tend to 'cut corners,' selectively reviewing only some files and missing others."
- "Position drift" -- reported issues "frequently don't match the actual code location."
- "Unstable quality" -- review quality "fluctuates significantly with minor prompt variations."
Its diagnosis: "a purely language-driven architecture lacks hard constraints on the review process." So steps that "must not go wrong" are ordinary code: precise file selection; "smart file bundling" that groups related files, each bundle running "as a sub-agent with isolated context"; rule matching by template engine rather than prompt; and separate modules that check comment positions and content. The agent keeps the parts that need judgement -- deciding what to look at and pulling in context -- with a toolset "distilled from deep analysis of tool-call traces in large-scale production data."
The trade-off, stated openly
On AACR-Bench -- Alibaba's own dataset of 200 real pull requests from 50 repositories in 10 languages, with 1,505 issues annotated by "80+ senior engineers" -- the README claims "significantly higher Precision and F1" than Claude Code on the same underlying model, "while consuming only ~1/9 of the tokens." It also says "its Recall is lower than general-purpose agents — a deliberate trade-off favoring precision over noise." These are self-reported results on a vendor-built benchmark.
Two practical details: ocr review --from main --to feature-branch reviews a branch since it diverged, and interrupted reviews can be resumed by session ID. A "delegation mode" lets your existing coding agent do the review, with OCR handling only file selection and rules, and needs no separate model configuration.
What developers are building
This is part of a wider pattern in this digest's coverage: wrapping agents in deterministic scaffolding wherever correctness matters. eHealth's fixed "static node" for regulated speech (October 9) and tester-army's replayable agent-written tests (October 6) take the same approach in other domains.
Analysis: the precision-over-recall choice is the right default for automated review in CI. Reviewers stop reading a bot that cries wolf, and a quiet, accurate reviewer gets used. Teams adopting it should measure what it misses on their own past pull requests before treating a clean review as a pass.
Alibaba's open-sourced review CLI puts file selection, bundling, rule matching and comment positioning in deterministic code and leaves only judgement to the agent, claiming higher precision at about a ninth of the tokens on its own benchmark, with lower recall by design.