What PRaline is
PRaline is an interactive terminal tool that reviews GitHub pull requests. It fetches a PR's diff and comment thread, asks Claude for a review, then walks you through each proposed comment so you can accept, reject or edit it before anything is posted under your account.
Three decisions shape the whole codebase. Keep them in mind before touching anything:
- Claude Code is the LLM backend. Reviews run through the
claudeCLI on the user's existing subscription. There is no Anthropic API key and no billed API call anywhere. - PRaline never writes code. The GitHub layer can read content and post comments or issues. It has no code path that pushes, merges, approves or creates branches.
- A human approves every comment, outside
praline auto. Nothing else reaches GitHub without going through the interactive approval loop first.
The whole tool is about 3,000 lines of Python across twelve small modules, with two runtime dependencies: requests and markdown.
Dev setup
You need Python 3.12+, uv, the claude CLI logged in, and a GitHub token in GITHUB_TOKEN or GH_TOKEN (Contents: read, Pull requests: read/write, Issues: read/write; a classic token with repo scope also works).
git clone git@github.com:ameroyer/PRaline.git
cd PRaline
uv sync
# run your working copy against any repo you have PRs in
uv run praline --dir ~/src/some-test-repo --model sonnet
Lint with uvx ruff check . (config in pyproject.toml: line length 100,
rules E, F, I; prompts.py is exempt from
E501 because it is prose, not code). There is no test suite yet, so testing is manual:
point --dir at a scratch repo with an open PR and exercise the menu.
A first test suite would be a very welcome contribution; github.py and
reviewer.py are the natural starting points since both are easy to mock.
Reviews post real comments under your GitHub account. Use a scratch repo while developing, and answer "n" at the posting confirmation if you only want to inspect the output.
Code layout
Each module has one job. The import graph is a shallow tree with cli.py at the root and verdict.py / term.py as dependency-free leaves; nothing imports cli.
| Module | Lines | Job |
|---|---|---|
cli.py | ~465 | Entry point (praline = praline.cli:main). Argument parsing and the main menu. All input() prompts live here or in reviewer.py. |
github.py | ~550 | Every GitHub call, REST and GraphQL, via requests (no gh CLI dependency), plus local git helpers and the per-run comment cache. This is the safety boundary; see Invariants. |
reviewer.py | ~510 | Assembles the review request, flattens Claude's reply into postable items, runs the approval loop, posts accepted comments, writes the review log and re-requests your review. |
prompts.py | ~385 | The system prompts (PR review, codebase scan, repo knowledge, PR history, module map) plus KB_STYLE, the writing rules shared by the knowledge-base prompts. The review output schema is defined here, in prose. |
memory.py | ~220 | Decides what the knowledge base says: scans the codebase, folds in merged PRs, guards against a rewrite erasing the previous document. Rendering lives in render.py. |
monitor.py | ~160 | The unattended watch loop (praline monitor): one loop around auto.run_auto that survives failures and waits out an exhausted budget. See Monitor mode. |
budget.py | ~155 | The rolling-window token cap. Pure, imports nothing from the package, and claude_client.ask is its only caller. See Monitor mode. |
auto.py | ~185 | Non-interactive mode (praline auto): picks PRs, reviews them, posts every comment, no prompts. See Auto mode. |
slack.py | ~330 | Optional Slack notifications: config loading, user resolution, the author/reviewer group chat, and the reviewer round-up. Depends only on verdict and term. |
verdict.py | ~90 | What a finished review is, as pure functions over dicts: the ready/minor/wip vocabulary, flatten_review, overview_of, count_items, reviewed_entry. Imports nothing, so any module can read a review without pulling in another. |
watch.py | ~205 | What changed since the last look (seen_state.json, the 🆕/🔄 icons) and order_prs, the stack-aware review order shared by the CLI and auto mode. |
claude_client.py | ~190 | Owns the model: ask() runs a single headless claude turn and returns its text, extract_json() parses a reply that was meant to be JSON. Holds the tool allow/deny lists and strips credentials from the child environment. |
config.py | ~195 | Paths and load/save helpers for the .praline/ files, plus Run, the dataclass holding one invocation's settings. No logic beyond file IO. |
render.py | ~150 | Turns the knowledge base into documents: the combined markdown file, and the HTML page with the stat tiles and module map. Calls nothing but markdown and the template, so it can be exercised on its own. |
hardness.py | ~155 | The four review depth levels. Mostly prose: one prompt addendum per level, plus clamp, label and explores. Imports nothing. See Review depth. |
graph.py | ~105 | Asks Claude for the repo's module map and validates what comes back, then labels the counts github.repo_counts returns. Runs no subprocess of its own. See The module map. |
mcp_server.py | ~370 | The optional MCP server (praline-mcp), exposing the same operations as tools for Claude. See The MCP server. |
term.py | ~30 | ANSI colors, _c / _rule, and confirm(), the shared yes/no prompt. Used by every module that prints. |
templates/ | ~540 | Not code: the standalone HTML shell for the knowledge base, filled in by render.knowledge_html. Includes the dependency-free JavaScript that lays out the module map. |
Two small conventions keep the plumbing out of the way. Settings for one invocation
(repo_dir, repo, model, reviewer_login,
request_review, slack) travel as a single
config.Run, built once in cli.main and handed to
_main_menu, _do_review and run_auto. A new setting
means one field, not five signatures. And a finished review is described in exactly one
place, verdict.reviewed_entry, so the CLI summary and the Slack round-up cannot
drift apart.
How a review runs
Menu option 1 ends up in cli._do_review, which drives this sequence. Note step 2: the knowledge base and the review log are appended to every review prompt, which is what makes reviews build on history rather than starting cold each time. auto and monitor warn when a repo has none rather than building one unattended.
The sequence:
github.list_open_prslists open PRs; the user picks one and sees a small status line (diff size, who has commented).reviewer.review_prfetches the raw diff and the full existing conversation (top-level comments plus line-level threads, tagged with comment ids), builds the system prompt fromprompts.DEFAULT_REVIEW_PROMPT(or a custom prompt file) plus the knowledge base, and makes oneclaude_client.askcall.- Claude must answer with a single JSON object: a
summary,repliesto existing threads (withreply_to_idand an optionalresolvedflag), newcomments(severitybug/warning/nit) andbugs.review_prstrips markdown fences defensively andjson.loadsthe rest. reviewer.flatten_reviewturns that object into one ordered list of postable items (summary, replies, bugs, comments). Both the approval loop and auto mode consume that shape, so ordering changes hit both.reviewer.run_approval_loopprints an overview of everything, then goes item by item: accept, reject or edit. The summary is the first item and, if accepted, becomes the top-level PR comment.- After a final confirmation,
reviewer.post_accepted_commentsposts each item: replies go into their original thread (and can mark it resolved), items with a file and line become line comments (pinned to the head sha the review was written against), the rest become general comments. - Then three tail steps, in this order and all non-fatal:
reviewer.log_reviewrecords the review in.praline/review_log.json,reviewer.rerequest_reviewputs you back on the PR's reviewer list, andslack.notify_reviewposts to Slack if--slackis on. Logging comes first on purpose: the memory must survive a Slack or GitHub hiccup in the other two.
Every reply carries a verdict alongside the summary: status is one of
ready / minor / wip, normalized by
verdict.status_key (which also accepts a full phrase like "needs minor
revisions") and rendered by verdict.status_label. It is shown at the top of the
approval overview, in the auto-mode summary and in Slack. An unusable value renders as
"❔ No status given" rather than being guessed at.
On a PR that already has a conversation, the prompt makes replying the primary job: the model must consider every existing thread before raising anything new, and an empty list of new comments is treated as a good outcome, not a failure.
Review depth
--hardness N (-H, default 0) says how hard to
look. It is one knob with four settings, and it lives entirely in
hardness.py: each level is a Level holding a name, a one-line
blurb for the UI, and an addendum that reviewer._build_review_prompt
appends to the system prompt.
| Level | Name | What changes |
|---|---|---|
| 0 | light | Default. Only what would change the author's mind or the shape of the code. At most five comments, no style nits. |
| 1 | standard | A full pass over every changed file. |
| 2 | thorough | Adds a checklist: edge cases, error paths, resource handling, invariants, security, tests, real performance problems. |
| 3 | exhaustive | Adversarial audit, and the review runs inside a read-only checkout so Claude can read around the diff. |
Two properties are worth preserving if you touch this. Depth is orthogonal to the prompt: it says how hard to look, not what to look for, so it is appended on top of a custom prompt file rather than replacing it. And the level never changes the response contract: the same JSON, the same severities, the same approval loop at every setting, so nothing downstream has to know which level ran.
Level 3 is the only one that changes machinery. hardness.explores(level) is
true from EXPLORE_FROM up, and reviewer._ask_in_checkout then
fetches the PR head, adds a temporary detached worktree, and runs the turn there with
Read/Glob/Grep and the credential denylist, the same
guardrails as the codebase scan. If the checkout cannot be made, it falls back to reviewing
the diff alone rather than failing: the checkout is a convenience, not the review.
Adding a level means adding one entry to LEVELS. The CLI help, the menu, the
MCP tool description and clamp() all read the dict, so nothing else needs
touching.
Suggested changes
When a fix is small and mechanical, the prompt asks for a
```suggestion block instead of a description of the fix. GitHub renders that
as a Commit suggestion button. This is prompt work plus one plumbing change: review
comments can now span a range.
comments[].start_lineis optional in the schema.github.post_pr_review_commentsendsstart_line/start_sideonly when it is strictly less thanline, because GitHub rejects anything else, so an equal or larger value silently degrades to a single-line comment rather than erroring.reviewer.has_suggestionis the one place that decides whether an item carries a suggestion. The approval loop uses it twice: to mark the item 🔧, and to print the body unwrapped, sincetextwrap.fillwould destroy the indentation GitHub commits.- The prompt is explicit that a wrong suggestion is worse than none. A committable-looking bad fix is the failure mode here, so anything uncertain is meant to come back as prose.
Auto mode
praline auto [PR ...] runs the same review as the interactive menu, minus
every prompt: auto.run_auto picks PRs, reviews each with
reviewer.review_pr, and posts every proposed item straight through
reviewer.post_accepted_comments. There is no approval loop.
- A PR qualifies if its number was passed as an argument, or if it is not a draft and either has never been reviewed here or has been commented on since it last was.
auto._has_new_activitychecks the PR's ownupdated_atfirst and only pays for the two comment endpoints when that says no. The fetch exists solely to catch a comment edited in place without the PR being bumped.- Per-user "last reviewed" timestamps live in
.praline/auto_state.json(config.load_auto_state/save_auto_state), keyed by PR number, and are only updated after a PR is actually processed. - Before reviewing anything, it refetches full stats per candidate (the list endpoint
omits
changed_files, same gotcha asget_pr_status) and sums them. Past--max-changed-filesit asks a single yes/no confirmation; this is the only prompt auto mode ever shows. - Each review is flattened by
reviewer.flatten_review, the same callrun_approval_loopmakes, and every item is posted unconditionally. - Each PR gets the same tail steps as an interactive review: log, re-request (unless
--no-request-review), notify. - A per-PR result (files changed, verdict, comments added, replies left, threads resolved,
or the error if the review call failed) is collected and printed as a summary table at the
end, and with
--slackDMed to you as a round-up grouped by verdict.
This is the one intentional exception to the "no auto-posting" invariant below: a human still approves every comment in the interactive menu, but auto mode exists precisely to skip that step for PRs the user has pointed it at.
What starts a review
select_prs takes on_new_commits, which picks between two
predicates. _has_new_comments is the default in both unattended modes: a PR
requalifies only when someone has commented. _has_new_activity is the wider
one, reached with --review-new-commits, where a push counts too.
The narrow one is the default because nobody asked for the review a push would trigger. An author iterating on a branch would collect a near-identical review per commit, which is noise for them and budget for the user. A comment is someone addressing the review, and that is worth answering.
Both predicates share _commented_since and differ only in how they use
pr.updated_at. For the wide one it is a cheap positive: anything that
bumps it counts, so the comment endpoints are only paid for when it says no. For the narrow
one it is a cheap negative: a PR nothing has touched cannot have been commented on,
but a bump alone proves nothing, so the comments are still read. Do not collapse them into
one shortcut; the asymmetry is the point.
Monitor mode & budgets
praline monitor is a loop around auto.run_auto, which already
picks the PRs worth reviewing. monitor.py adds only what turns a one-shot
command into something you can leave running. Keep that division: selection and review
logic belongs in auto, not in the loop.
What requalifies a PR is decided in auto.select_prs and shared by both modes:
see Auto mode.
The one change monitor mode forced on auto is
run_auto(..., confirm_budget=False). The changed-file check is the only
interactive moment in that path; unattended, a batch over the limit is skipped and left for
a human rather than confirmed by default. run_auto also returns its per-PR
results now, which is what the loop reports on.
Each pass empties the comment cache first
(github.forget_all_comments). That cache is invalidated on a post and
otherwise lives for the process, which is right for a command that exits and wrong for one
that looks a hundred times: without clearing it, a comment left between two passes would
never be seen and a re-review would be written against a conversation that had moved on. If
you add long-running behaviour anywhere else, it needs the same treatment.
The knowledge base is refreshed when PRs merge while watching, via
memory.merged_since_last_update. That refresh redraws the module map only when
the repo has never had one: the map is the most expensive call PRaline makes and codebase
shape moves far more slowly than merged PRs, so redrawing an existing one unattended would
spend most of the budget on the part that changed least. A repo only ever watched would
otherwise never get a map at all, which is why the first one is exempt.
The page is re-rendered at startup by memory.rerender_html,
which reads the documents already on disk and calls nothing. The page and the documents age
at different rates: documents change with the repo, the page changes with PRaline's
template. Without this, a knowledge base built before a template change keeps producing the
old page, and the symptom is a section the code plainly emits being absent from the file
the user is looking at.
The budget
A rolling window, enforced before a call, not after.
budget.guard() runs at the top of claude_client.ask, so an
exhausted allowance means a review never starts, rather than one being abandoned
half-posted. auto checks it once per PR and breaks out cleanly, leaving the
rest to requalify next run.
Two details that look like details and are not. Cached reads are excluded from the
count (budget.billable): a cached system prompt reports tens of
thousands of read tokens on a trivial call, at a fraction of the price, so counting them
would exhaust a sensible cap in four calls and describe nothing real. And
spending is persisted to .praline/budget.json: a monitor that
crashes and restarts must not receive a fresh allowance, which is exactly when a cap
matters. A stored window that no longer matches the requested one is discarded rather than
reinterpreted.
DEFAULT_TOKENS_PER_HOUR is 40,000, and monitor mode is the only mode that gets
a budget without being asked, because it is the only one that runs unattended for hours. The number
is measured, not guessed: one review of a small PR at depth 0 costs about 12.3k billable
tokens, so the default is roughly three an hour. If you change the default, re-measure
rather than reasoning about it; --tokens-per-hour takes an explicit
0 to disable.
The active budget is module state (budget.ACTIVE) rather than a
config.Run field, because claude_client.ask is called from a dozen
places that have no business knowing about budgets. Default is None, and then
guard() and record() do nothing at all.
What's new, and PR order
watch.py answers two questions that have nothing to do with the model.
- What moved since last time.
classifytags each open PRNEW(absent from.praline/seen_state.json),UPDATED(itsupdated_atmoved) orSEEN, andrun_checkprints the digest behindpraline check(--no-markto report without acknowledging) and menu entry 4. Only those two write state (mark_seen); the interactive PR list reads it and never acknowledges, so the icons stay put until you say so. First run has nothing to compare against, so everything isNEWexactly once.is_first_lookkeys offlast_checked, not the PR map, because a repo with no open PRs stores an empty one. - What order to review in.
order_prssorts by PR number ascending, except that a PR whose base branch is another open PR's head branch is a stack and comes after its parent. Stacks are walked depth first so they stay contiguous. A cycle or a parent that isn't open falls back to number order; nothing is ever dropped.stack_parentsbuilds the whole parent map in one pass and feeds the↳ on #5markers in the list.
Both the CLI picker and auto.select_prs go through this module, so ordering and
icons cannot drift apart between the two modes.
Slack notifications
Entirely optional and off unless --slack is passed. cli._load_slack
resolves the config once at start-up and exits if it is broken; after that every Slack failure
is printed and swallowed, because the review it would report on is already posted.
- Config.
slack.load_configsearches$PRALINE_SLACK_CONFIG, then~/.config/praline/slack.json, then<repo>/.praline/slack.json. The token may come from$SLACK_BOT_TOKEN/$PRALINE_SLACK_BOT_TOKENinstead, so it need not touch a file at all. Nothing here ever writes a token. - Who gets the message.
notify_reviewposts to the PR author and the reviewer (cfg.reviewer_login, the GitHub login PRaline authenticated as). Two recipients meansconversations.openand a group DM, which is the reasonmpim:writeis required; one means a directchat.postMessageto the member ID, which needs no extra scope. Unmapped people are skipped, not fatal. - Resolution is cached. A mapping value may be a member ID, an
@handleor an email; handle lookups page throughusers.list, so results are memoized on the config object for the run. - The round-up.
notify_digestDMs the reviewer a list of everything just reviewed, grouped by verdict, ready-to-approve first. Sent at the end ofrun_auto, and on quitting an interactive session in which anything was reviewed.notify_digestitself returns early on an empty list, so callers do not need to guard the count.
The review log
.praline/review_log.json is PRaline's memory of its own work: one entry per
posted review with the PR number, title, author, head sha, branches, verdict, the overview
text and the comment counts, capped at the last 200 by
config.append_review_log.
reviewer._format_review_log feeds the newest entries back into later review
prompts, always including earlier passes over the same PR and labelling them. That
is what makes a re-review build on the last one instead of repeating it, and what lets a
stacked PR be read in light of the PR below it, which was reviewed minutes earlier. It is
distinct from the knowledge base (lessons from merged PRs) and from
seen_state.json (what is new since you last looked).
Huge PRs
GitHub's API refuses to render a PR past roughly 300 changed files as a single diff:
get_pr_diff gets a 406 Not Acceptable and raises
DiffTooLargeError. reviewer.review_pr catches it and hands off to
_review_huge_pr, which warns the user ("huge PR, careful") and recovers in two
stages:
- Local diff.
github.fetch_pr_refsfetches the base branch and the PR's hidden head ref (pull/N/head), thenget_pr_diff_localproduces the same unified diff withgit diff origin/<base>...<head_sha>, no size limit. The PR's true size (fromgit diff --stat, sincePRInfocounts are 0 off the list endpoint) is printed before going further. - Explore mode. Interactively, the user is asked; in auto mode it is
enabled directly, since auto mode has no prompts.
github.add_pr_worktreechecks the PR head out detached in a temporary worktree, the diff is written toPRALINE_DIFF.patchat its root, and the review call runs withreadableset to the worktree and the scan's longer timeout.prompts.HUGE_PR_EXPLORE_ADDENDUMtells the model to page through the diff in slices and use the checkout for context. The worktree is force-removed in afinally. Declining explore mode inlines the full local diff instead, which works whenever it fits the model's context.
Either way the response contract is untouched: the same JSON schema comes back, and the
approval loop or auto-posting proceeds as for any other PR. The feature holds the four
holy principles (all six of them): MINIMALISM, nothing changes before a
406 and the tooling is the scan's, reused; CLEANLINESS, one prompt, one
schema, only the diff's transport changes; MODULARITY, git plumbing in
github.py, flow in reviewer.py, instructions in
prompts.py; EFFICIENCY, the diff stat steers the model's
reading budget toward the files that matter; SAFETY, the worktree is
temporary and detached, so the user's checkout, index and branches are never touched;
SECURITY, explore mode runs read-only with the credential denylist,
no shell and no network, exactly like the codebase scan.
The knowledge base
Menu option 2 calls memory.update_knowledge, which returns the
(markdown, html) paths so the CLI can print them. It writes into
.praline/ inside the reviewed repo, not inside PRaline:
| File | Written by | Contents |
|---|---|---|
repo.md | scan_codebase or build_repo_knowledge | Architecture, module reference, conventions, invariants, pain points. |
pr_history.md | build_pr_history | Lessons from merged PRs in the chosen window, every claim cited as (#123) and linkified to the PR page. |
knowledge.md | render_knowledge_md | The two halves stitched together under one title, for reading and grepping. |
knowledge.html | render_knowledge_html | The same document through templates/knowledge.html, standalone and publishable. |
knowledge_state.json | update_knowledge | Build timestamp, used to detect newly merged PRs. |
artifact_url.txt | you, optionally | Pointer to a published copy of the HTML. |
repo.md is built two different ways, and which one runs is decided by whether
it already has content:
- First build:
scan_codebase. A tool-enabled Claude turn runs withcwdset to the repo and reads it directly (Read,Glob,Grep). It follows entry points and imports rather than guessing from a file list, so this is the pass that produces the module reference and the invariants. Slow, and it runs once. Timeout isEXPLORE_TIMEOUT_S, not the default. - Every later build:
build_repo_knowledge. No tools. It looks up the default branch via the GitHub API, runsgithub.fetch_remote_branch, and readsgit ls-tree/git logoff the resultingorigin/<branch>ref instead ofHEAD, so a stale local checkout cannot skew it.git fetchonly updatesrefs/remotes/*, never the working tree, the index, or local branches. If your checkout was behind, a one-line note reports by how many commits; if the fetch failed outright, it says why and falls back to readingHEAD.
Fetching goes through one function, github.fetch_refs, and it tries twice.
First git fetch origin, using whatever credentials the remote is configured
for. If that fails it retries over HTTPS with the GitHub token PRaline already requires.
That second attempt is not a nicety: the remote is usually ssh, and an ssh agent is the one
thing PRaline cannot assume is there: a tmux session outliving the agent that started it,
or a machine with no key loaded, used to leave every fetch dead behind a bare "could not
fetch origin/main" while a perfectly good token sat in the environment.
The token reaches git through the environment and is referenced by name in the
credential helper, so its value never appears in argv where ps would show it to
every user on the machine, and never lands in .git/config. An empty
credential.helper= is passed first to reset the helper list, so a system helper
configured elsewhere cannot answer ahead of it with the wrong account. Both attempts use the
same explicit refspecs, so what ends up in refs/remotes/ does not depend on
which one won. Only a failure of both raises FetchError, carrying the first
line of git's own message, run through redact_url, since that message may
quote the remote.
Both markdown halves are fed into the review prompt on every review, which is the whole point: reviews get repo-specific over time.
Updates are edit passes, not rewrites. The previous document is included in the
prompt with instructions to merge, and _guard_against_erasure refuses any
result shorter than half the previous version, keeping the old file instead.
A .bak copy is written before each save. If you touch this flow, preserve
both protections; silently losing accumulated knowledge is the worst failure
mode this module has.
Staleness is checked on the way into a review. memory.last_updated_at reads
the build timestamp (falling back to repo.md's mtime for knowledge bases
built before that file existed), merged_since_last_update asks GitHub what
has merged since, and cli._offer_update_for_new_prs lists those PRs and
offers to refresh. A failure here prints a warning and gets out of the way; it must never
block the review the user actually asked for.
The module map
knowledge.html ends with a "Repo at a glance" section: stat tiles over a
node-link diagram of the modules. Three pieces, deliberately separated.
- Getting it.
graph.buildruns one read-only Claude turn against the checkout withprompts.ARCH_GRAPH_PROMPT, which asks for{nodes, edges}and nothing else. - Trusting none of it.
graph.normalizedrops what cannot be drawn: unknown node kinds fall back tomodule, duplicate ids, self-loops, repeated pairs, and any edge naming a node that was never declared. It never raises. A diagram missing two boxes is worth showing; a knowledge base that failed to build because the map came back malformed is not.memory.build_graphtakes the same line one level up: on any failure it keeps the previous map and carries on. - Drawing it.
render.pyemits the section and embeds the graph in a<script type="application/json">block; the template's JavaScript lays it out. No CDN, no build step, so the page stays a single self-contained file you can publish as an artifact.
The layout is a layered DAG, not a force simulation, so the same graph always draws the same picture. Cycles are broken before layering, with a depth-first walk that marks the edges closing a cycle; those are excluded from the layer assignment and rendered dashed. Skipping that step is not cosmetic: relaxing longest-path across a cycle pushes its nodes a layer deeper on every pass, and four mutually-importing modules end up strung across a canvas wide enough for thirteen.
The stat tiles come from github.repo_counts, never from the model. That function lives in github.py because every git call does; graph.stats only attaches labels. A number
a model wrote down is a number nobody can check. Tiles that come back zero are dropped, so
a shallow clone renders fewer tiles instead of a row of noughts.
Node kind is carried by shape and outline (a ring and ▸ for an entry
point, dashed for external), not by color alone, and a table view sits under the diagram.
If you restyle it, keep both: the theme's pastels are decorative and do not survive a
colorblind-separation check as a categorical encoding.
The MCP server
praline-mcp exposes the same operations as MCP tools, so Claude can drive
PRaline directly. It is an optional extra (uv sync --extra mcp); nothing else
in the package imports mcp_server, so the dependency stays optional in fact and
not only in the metadata.
Each tool resolves its repo the same way, through _resolve: the explicit
repo_dir argument, else PRALINE_REPO_DIR, else the process's
working directory, which is the repo you are in when Claude Code launches the server.
Two constraints shape the module, and both are easy to break by accident:
-
Drafting and posting are separate tools.
review_pr_draftnever writes to GitHub; it stashes the result in_DRAFTSand returns the numbered draft.post_reviewposts the indices the user approved, out of that stash rather than by re-running the review, which would come back subtly different. A model that reads "review this PR" as "review and post it" therefore cannot post anything. -
stdout belongs to the protocol. The rest of PRaline narrates as it works,
and on a stdio transport one stray
printcorrupts the JSON-RPC stream. Every tool body runs inside_captured(), and_reply()hands the captured text back as the result'slogso the progress is not simply lost. If you add a tool, wrap it the same way.
The Claude backend
claude_client.ask() is the only place PRaline talks to a model. It runs:
claude -p --output-format json --model <model> \
--allowedTools <scoped-read-rules> --disallowedTools <secret-globs> \
--setting-sources "" --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
--name <session name> --system-prompt-file <temp file>
Details worth knowing before you modify it:
- Nothing that grows with the repo goes into argv. Linux caps a single
argv entry at 128 KiB (
MAX_ARG_STRLEN), whatevergetconf ARG_MAXreports, and both of the big strings pass that on a real repo. The user message (typically a large diff) is piped through stdin. The system prompt carries the knowledge base, the PR history lessons and the review log, so it grows with the repo; it goes through a 0600 temp file written by_write_system_promptand deleted in afinally. Passing either as an argument fails with[Errno 7] Argument list too long, and it fails later the smaller your knowledge base is, which is why it first showed up in monitor mode. If you add another argument, keep it bounded. - Every turn is named, and the name is required.
session_nameis what the session is called in/resume, the prompt box and the terminal title, instead of a hex id. Callers pass the tail only;ask()prependsSESSION_PREFIX, so one grep finds every session PRaline started and no call site can spell the prefix differently. Reviews getPRaline: owner/repo PR #12 (review 2)fromreviewer._session_name(), which counts earlier passes over that PR in the review log; the knowledge base, PR history, codebase scan and module map name themselves. Build a name from what PRaline already knows, never from PR text: a name goes into argv, where the diff deliberately does not, and a PR title is written by whoever opened the PR. - There are exactly two call shapes, and the confinement is structural.
ask()takesreadable, one directory, and there is notools=parameter any more. Omit it, as reviews and update passes do, and the turn gets no tools. Pass it, as the codebase scan, the module map and explore mode do, and it getsRead,GlobandGrepscoped to that directory by_read_only_rules, plus the secret denylist and_ISOLATION_ARGS, whichask()applies itself so no call site can omit them. A bareReadallows reading any path on the machine, which is why the old signature was removed rather than documented. Scoping every read tool matters:Greptakes its ownpathand returns the matching lines of a file it was never allowed to open. Nothing else should need tools; if you think it does, see Secrets and data flow first. - No settings, no MCP servers. Explore mode runs with the working directory
set to a checkout of the PR under review, so a
.claude/settings.jsonthere would install hooks and a.mcp.jsonwould launch a server command: arbitrary code execution triggered by opening a pull request._ISOLATION_ARGScloses both, and keeps a review away from the user's own MCP servers. Never restore a settings source to widen a permission; scope the rule instead. - Enforcement is the permission layer, not the model. Verify by asking a
turn to read outside its directory and finding the attempt in the reply's
permission_denials. A model declining is not evidence of anything. - Every failure mode (timeout, non-zero exit, empty stdout, bad JSON, error payload, empty result) raises a
RuntimeErrorcarrying truncated stderr/stdout. Keep that style: errors here surface directly in the CLI and are the user's only debugging signal. - The default timeout is 600 seconds; big diffs on big models are slow. The scan gets 2,400.
- The subprocess runs with
env=_clean_env(): the current environment minusGITHUB_TOKEN,GH_TOKEN,SLACK_BOT_TOKENandPRALINE_SLACK_BOT_TOKEN. The turn has no use for them, and a process that never receives a secret cannot leak it, whether through a tool, a hook, an MCP server, or a diff written to talk the model into echoing its environment. Add any new credential variable to that tuple the moment you introduce it.
Secrets and data flow
The knowledge base is the one place where reading and publishing meet. Whatever lands in
repo.md is fed back into every later review prompt, can be posted to GitHub as
part of a comment, and may be published as an artifact. So the scan is constrained on two
levels, and both matter:
- The CLI blocks credential files. The secret denylist in
claude_client.pyis a list ofRead(<glob>)deny rules (.env*,*.pem,*.key,id_rsa*,.netrc,secrets.*,credentials.*and more). They match credential files on purpose: a blanket*token*would hidetokenizer.py, and a scan that quietly skips real code is its own failure. Deny rules override the allowlist and apply toGrepas well asRead, so a blocked file cannot be laundered through a content search. This is enforced by the harness, not by the model's judgement. - The prompt handles what globs cannot. A key hardcoded in ordinary
source is not matched by any denylist, so
INIT_CODEBASE_PROMPTrequires recording that such a secret exists and where, never its value.
If you extend the scan, add the glob before adding the capability. Other things that hold:
- The GitHub token is read from the environment in
github._tokenper request and only ever sent as anAuthorizationheader toapi.github.com. It is never logged, never written to.praline/, never interpolated into a message, and never put in a URL. Error paths raise on the response, which does not carry request headers. Same rules for the Slack token, which only reachesslack.com. - A token can arrive from somewhere PRaline does not control: a git remote of the form
https://ghp_xxx@github.com/owner/repo.github.redact_urlstrips the userinfo part before that string can reach an error message. If you print a remote URL anywhere, print it through that function. - Credentials are stripped from the review subprocess environment; see
The Claude backend. The list in
claude_client._CREDENTIAL_ENV_VARSincludesFETCH_TOKEN_ENV, the variablegithub.fetch_refspasses the token to git in. It is defined next to the strip list and imported from there, because naming it in two places is how it ends up stripped in only one. cli._warn_if_praline_dir_committablerunsgit check-ignoreat start-up and warns when.praline/is not ignored in the target repo. That directory holds the knowledge base, the review log and possibly a Slack config.- System prompts are passed on argv, so they are visible to other users on the same machine via the process list. That covers the knowledge base and the diff-free part of a review, which is repo-internal but not secret. Do not move anything sensitive into that argument.
.praline/lives in the reviewed repo and describes it in detail. PRaline warns when it is not ignored but cannot fix it, since it never commits.- A PR diff or description is untrusted input that ends up in a prompt, in the same
prompt as the knowledge base, the PR-history lessons and the review log. The interactive
approval loop is what stands between that and a posted comment.
autoandmonitorremove it deliberately, so on those paths a PR crafted to make the model quote its own context can put internal notes into a public comment. That is the residual risk of unattended review and the reason both are opt-in.
Invariants
These are promises the README makes to users. A PR that breaks one will be rejected regardless of how useful the feature is.
- No write-to-code endpoints. Every function in
github.pymaps to a read endpoint, a comment endpoint, or the review-request endpoint. When adding a GitHub call, keep it inside the ALLOWED list documented at the top of that file, and note what is not there: approving, merging, and opening or editing issues. - No auto-posting outside
praline auto. The interactive menu's comments flow throughrun_approval_loopand the final confirmation; nothing incli._do_reviewmay skip them.auto.run_autois the one deliberate exception, gated behind its own subcommand. - No API key. The model backend stays the
claudeCLI. Do not introduceanthropicSDK calls or key handling. - No tools beyond read-only, and never unscoped. A tool-enabled call names its directory through
readableand gets nothing else. A write or shell tool would break the "never touches code" promise outright, and an unscoped read tool hands a prompt-injected PR the whole filesystem. - Credentials never reach a subprocess or a log line. Tokens are read from the environment at the point of use, sent only to their own API host, and stripped from the
claudechild environment. Nothing writes one to disk, and no error message may interpolate one. - Slack only ever sends, and only to people. The four methods called are
auth.test, a user lookup,conversations.openandchat.postMessageto a DM or group DM. No channels, no history reads, and a Slack failure never rolls back or blocks a review. - Nothing posts from the MCP server without a second call.
review_pr_draftdrafts and returns;post_reviewpublishes what the user picked. Any new tool that writes to GitHub stays separate from the tool that produces the content, for the same reason the interactive loop exists. - Every review is grounded in the accumulated knowledge.
reviewer._build_review_promptis the single place that appendsrepo.md,pr_history.mdand the review log, and every review path goes through it, including huge-PR explore mode, depth 3, and the MCP tool. A new path that assembles a prompt some other way silently loses the repo-specific half of the review. The module map is deliberately not in there; it is for the HTML page. - The rendered knowledge base never carries live markup.
render._NoRawHtmlderegisters Markdown'shtml_blockpreprocessor andhtmlinline pattern, so raw HTML inrepo.mdorpr_history.mdbecomes visible text. Both files are model output over content that includes attacker-controlled PR titles and bodies, and the page they produce is opened overfile://and published as an artifact. Do not "fix" this by escaping the HTML stash instead:fenced_codeshares that stash, and every code block in the document would be double-escaped. - Anything untrusted printed to the terminal goes through
term.plain. PR titles, author names and comment bodies are written by other people, and GitHub does not strip control bytes on the way out. A title carrying\r\x1b[2Kerases its own line and reprints whatever the author chose, which in a tool driven by[a]ccept/[r]ejectand[Y/n]prompts is a way to forge what the user thinks they are approving. Colour is applied by_caround the sanitized text, never taken from the text. If you add a print of anything that came from GitHub or from the model, wrap it. - Anything quoted into a Slack message goes through
slack._escape. Slack reads<...>as markup, so an unescaped PR title can forge a link or ping the room. The<url|label>syntax this module builds is applied after escaping, never to it. - Unattended modes stay capped and non-draft. Monitor mode gets a token budget whether or not one was asked for, and drafts never qualify for an automatic review. Both are what makes leaving it running defensible.
- The review reply is JSON with a fixed schema. The schema lives in
prompts.DEFAULT_REVIEW_PROMPTand its consumers arerun_approval_loopandpost_accepted_comments. Change all of them together, and remember users can supply custom prompt files that must still produce the same schema.
Where changes go
| You want to... | Touch |
|---|---|
| Add a menu action | cli._main_menu plus a _do_* helper next to the existing ones; take a config.Run rather than loose arguments. |
| Add a global setting or flag | cli._add_common_args for the flag, a field on config.Run for the value. Both modes then see it without any signature changes. |
| Change how reviews behave | prompts.py first. Most behavior (tone, priorities, reply-before-comment) is prompt-defined, not code-defined. |
| Add a GitHub operation | github.py, respecting the ALLOWED list. Return plain data (dataclasses or dicts); rendering belongs to callers. |
| Change what a review sees | reviewer.review_pr and _format_existing_conversation for the user message, _build_review_prompt for the system side. |
| Change the approval UX | reviewer.run_approval_loop, _display_comment, _prompt_action. |
| Change knowledge-base content | prompts.INIT_CODEBASE_PROMPT (first build), INIT_REPO_PROMPT (updates) or INIT_PR_HISTORY_PROMPT; the surrounding pipeline is memory.py. |
| Change how the knowledge base reads | prompts.KB_STYLE. It is appended to all three knowledge-base prompts, so the rules stay in one place. |
| Restyle the knowledge-base page | praline/templates/knowledge.html for the shell and the map's JavaScript, render.py for the fragments Python builds. |
| Add or change a review depth | hardness.LEVELS, and nothing else. The CLI help, the menu, the MCP tool description and clamp() all read the dict. |
| Change what the module map shows | prompts.ARCH_GRAPH_PROMPT for what Claude is asked, graph.normalize for what survives, the template's script for how it draws. |
| Add an MCP tool | mcp_server.py: a @server.tool() function wrapping its body in _captured() and returning _reply(...). Keep anything that writes to GitHub behind its own explicit tool. |
Add a stored file under .praline/ | config.py for the path and load/save helpers, then use them from wherever needs it. |
| Change auto-mode selection or posting | auto.select_prs for which PRs qualify, auto.run_auto for the review/post/summary loop. Monitor mode inherits both. |
| Change what monitor mode does per pass | monitor.run_monitor for the loop, monitor._report for what a pass prints. Anything about which PRs qualify belongs in auto. |
| Change the review order or the new/updated icons | watch.order_prs and watch.classify. Both modes read them, so a change lands everywhere at once. |
| Change a Slack message | slack.review_message (per PR) or slack.digest_message (the round-up). Delivery is notify_review / notify_digest; keep failures non-fatal. |
| Add or change a verdict | verdict.STATUS_LABEL and verdict.status_key, plus the status guide in prompts.DEFAULT_REVIEW_PROMPT. The three renderers (CLI, auto summary, Slack) all read the same helpers. |
| Record something about past reviews | reviewer.log_review for the entry, _format_review_log for what the prompt sees. |
Gotchas
- A model reply that should be JSON goes through
claude_client.extract_json. It tolerates a fenced block or prose either side of the object, and it is shared by the review parser and the module-map parser. Do not write a third one. - The module map is laid out from a cycle-free graph.
graph.normalizeremoves self-loops but not cycles; the template breaks those at layout time. If you rewrite the layering, keep the depth-first cycle break or a mutually-importing package will draw itself across a canvas ten times too wide. - A suggestion block must not be re-wrapped. Its indentation is what GitHub commits, so
_display_commentprints those bodies verbatim. Any new renderer of comment bodies has to do the same. - Resolving review threads is GraphQL-only. GitHub's REST API cannot mark a conversation resolved, so
github.pycarries two small GraphQL calls (find_review_thread_id,resolve_review_thread). Everything else is REST. reply_to_idmust be the thread's root comment. Replying to a reply detaches the comment from its thread. The prompt instructs the model accordingly; keep that instruction if you rewrite it.get_merged_prspaginates byupdated, notmerged_at. The early-exit condition is deliberately conservative (a whole page must be older than the cutoff) and there is a hard cap of 10 pages. Easy to break if you "simplify" it.- Claude sometimes wraps JSON in code fences.
reviewer._extract_jsontries the raw text, then a fenced block, then the outermost{...}, and keeps the first that parses. Removing this "redundant" code will produce intermittent parse failures. - The same goes for prose around a document.
memory._strip_preambledrops anything before the first markdown heading, because a scan that used tools sometimes opens with a sentence about what it read. The prompts forbid it; the strip is the backstop. - The HTML template uses
safe_substitute, notsubstitute. It contains literal$characters in its prose, whichsubstitutewould reject as bad placeholders. Keep the placeholders ($title,$repo,$toc,$body) distinctive if you edit it. - The knowledge files live in the reviewed repo. Consider suggesting
.praline/for that repo's.gitignore; PRaline itself never commits anything, so it cannot do this for the user. get_repo_structure/get_recent_commitstake an explicit ref. They readgit ls-tree/git log <ref>, not the working tree orHEAD. Callers outsidebuild_repo_knowledgemust still pass a ref (e.g."HEAD") if they want the local checkout.- Subparser defaults must be
SUPPRESS._add_common_argsis called on the top-level parser and on every subparser so--dirworks on either side of the subcommand. argparse applies a subparser's defaults after the top-level parse, so a real default on the subparser silently overwrites what the user passed before the subcommand, which is why_add_common_args(..., subcommand=True)swaps every default forargparse.SUPPRESS. Add a new shared flag through that function, never directly to a subparser. - Comment lists are cached per run.
github._COMMENT_CACHEmakes the three readers of a PR's comments (status line, prompt conversation, reply lookup) cost one fetch each instead of three. Every posting function callsforget_pr_comments; if you add another write path, call it too or the next read will be stale. - Line comments take the head sha from the caller.
post_pr_review_commentno longer refetches the PR per comment. Passpr.head_sha, which also pins the comment to the commit that was actually reviewed. - You cannot request a review from the PR author. GitHub answers 422, so
rerequest_reviewskips your own PRs. Anything else it fails on is printed and ignored. PRInfocounts are only populated on single-PR fetches. The list endpoint omitsadditions/deletions/changed_files, so those default to 0 untilget_pris called; that is whyget_pr_statusandauto.run_autoboth refetch.draftandupdated_at, by contrast, are present on the list endpoint already.