$ cat developer_guide.md

🍫 PRaline Developer Guide

Everything you need to start hacking on the code review chocolate.

What PRaline is

PRaline is an interactive terminal tool that reviews GitHub pull requests. It fetches a PR's diff and comment thread, asks Claude for a review, then walks you through each proposed comment so you can accept, reject or edit it before anything is posted under your account.

Three decisions shape the whole codebase. Keep them in mind before touching anything:

  • Claude Code is the LLM backend. Reviews run through the claude CLI on the user's existing subscription. There is no Anthropic API key and no billed API call anywhere.
  • PRaline never writes code. The GitHub layer can read content and post comments or issues. It has no code path that pushes, merges, approves or creates branches.
  • A human approves every comment, outside praline auto. Nothing else reaches GitHub without going through the interactive approval loop first.

The whole tool is about 3,000 lines of Python across twelve small modules, with two runtime dependencies: requests and markdown.

Dev setup

You need Python 3.12+, uv, the claude CLI logged in, and a GitHub token in GITHUB_TOKEN or GH_TOKEN (Contents: read, Pull requests: read/write, Issues: read/write; a classic token with repo scope also works).

git clone git@github.com:ameroyer/PRaline.git
cd PRaline
uv sync

# run your working copy against any repo you have PRs in
uv run praline --dir ~/src/some-test-repo --model sonnet

Lint with uvx ruff check . (config in pyproject.toml: line length 100, rules E, F, I; prompts.py is exempt from E501 because it is prose, not code). There is no test suite yet, so testing is manual: point --dir at a scratch repo with an open PR and exercise the menu. A first test suite would be a very welcome contribution; github.py and reviewer.py are the natural starting points since both are easy to mock.

Reviews post real comments under your GitHub account. Use a scratch repo while developing, and answer "n" at the posting confirmation if you only want to inspect the output.

Code layout

Each module has one job. The import graph is a shallow tree with cli.py at the root and verdict.py / term.py as dependency-free leaves; nothing imports cli.

ModuleLinesJob
cli.py~465Entry point (praline = praline.cli:main). Argument parsing and the main menu. All input() prompts live here or in reviewer.py.
github.py~550Every GitHub call, REST and GraphQL, via requests (no gh CLI dependency), plus local git helpers and the per-run comment cache. This is the safety boundary; see Invariants.
reviewer.py~510Assembles the review request, flattens Claude's reply into postable items, runs the approval loop, posts accepted comments, writes the review log and re-requests your review.
prompts.py~385The system prompts (PR review, codebase scan, repo knowledge, PR history, module map) plus KB_STYLE, the writing rules shared by the knowledge-base prompts. The review output schema is defined here, in prose.
memory.py~220Decides what the knowledge base says: scans the codebase, folds in merged PRs, guards against a rewrite erasing the previous document. Rendering lives in render.py.
monitor.py~160The unattended watch loop (praline monitor): one loop around auto.run_auto that survives failures and waits out an exhausted budget. See Monitor mode.
budget.py~155The rolling-window token cap. Pure, imports nothing from the package, and claude_client.ask is its only caller. See Monitor mode.
auto.py~185Non-interactive mode (praline auto): picks PRs, reviews them, posts every comment, no prompts. See Auto mode.
slack.py~330Optional Slack notifications: config loading, user resolution, the author/reviewer group chat, and the reviewer round-up. Depends only on verdict and term.
verdict.py~90What a finished review is, as pure functions over dicts: the ready/minor/wip vocabulary, flatten_review, overview_of, count_items, reviewed_entry. Imports nothing, so any module can read a review without pulling in another.
watch.py~205What changed since the last look (seen_state.json, the 🆕/🔄 icons) and order_prs, the stack-aware review order shared by the CLI and auto mode.
claude_client.py~190Owns the model: ask() runs a single headless claude turn and returns its text, extract_json() parses a reply that was meant to be JSON. Holds the tool allow/deny lists and strips credentials from the child environment.
config.py~195Paths and load/save helpers for the .praline/ files, plus Run, the dataclass holding one invocation's settings. No logic beyond file IO.
render.py~150Turns the knowledge base into documents: the combined markdown file, and the HTML page with the stat tiles and module map. Calls nothing but markdown and the template, so it can be exercised on its own.
hardness.py~155The four review depth levels. Mostly prose: one prompt addendum per level, plus clamp, label and explores. Imports nothing. See Review depth.
graph.py~105Asks Claude for the repo's module map and validates what comes back, then labels the counts github.repo_counts returns. Runs no subprocess of its own. See The module map.
mcp_server.py~370The optional MCP server (praline-mcp), exposing the same operations as tools for Claude. See The MCP server.
term.py~30ANSI colors, _c / _rule, and confirm(), the shared yes/no prompt. Used by every module that prints.
templates/knowledge.html~540Not code: the standalone HTML shell for the knowledge base, filled in by render.knowledge_html. Includes the dependency-free JavaScript that lays out the module map.

Two small conventions keep the plumbing out of the way. Settings for one invocation (repo_dir, repo, model, reviewer_login, request_review, slack) travel as a single config.Run, built once in cli.main and handed to _main_menu, _do_review and run_auto. A new setting means one field, not five signatures. And a finished review is described in exactly one place, verdict.reviewed_entry, so the CLI summary and the Slack round-up cannot drift apart.

How a review runs

Menu option 1 ends up in cli._do_review, which drives this sequence. Note step 2: the knowledge base and the review log are appended to every review prompt, which is what makes reviews build on history rather than starting cold each time. auto and monitor warn when a repo has none rather than building one unattended.

The sequence:

  1. github.list_open_prs lists open PRs; the user picks one and sees a small status line (diff size, who has commented).
  2. reviewer.review_pr fetches the raw diff and the full existing conversation (top-level comments plus line-level threads, tagged with comment ids), builds the system prompt from prompts.DEFAULT_REVIEW_PROMPT (or a custom prompt file) plus the knowledge base, and makes one claude_client.ask call.
  3. Claude must answer with a single JSON object: a summary, replies to existing threads (with reply_to_id and an optional resolved flag), new comments (severity bug / warning / nit) and bugs. review_pr strips markdown fences defensively and json.loads the rest.
  4. reviewer.flatten_review turns that object into one ordered list of postable items (summary, replies, bugs, comments). Both the approval loop and auto mode consume that shape, so ordering changes hit both.
  5. reviewer.run_approval_loop prints an overview of everything, then goes item by item: accept, reject or edit. The summary is the first item and, if accepted, becomes the top-level PR comment.
  6. After a final confirmation, reviewer.post_accepted_comments posts each item: replies go into their original thread (and can mark it resolved), items with a file and line become line comments (pinned to the head sha the review was written against), the rest become general comments.
  7. Then three tail steps, in this order and all non-fatal: reviewer.log_review records the review in .praline/review_log.json, reviewer.rerequest_review puts you back on the PR's reviewer list, and slack.notify_review posts to Slack if --slack is on. Logging comes first on purpose: the memory must survive a Slack or GitHub hiccup in the other two.

Every reply carries a verdict alongside the summary: status is one of ready / minor / wip, normalized by verdict.status_key (which also accepts a full phrase like "needs minor revisions") and rendered by verdict.status_label. It is shown at the top of the approval overview, in the auto-mode summary and in Slack. An unusable value renders as "❔ No status given" rather than being guessed at.

On a PR that already has a conversation, the prompt makes replying the primary job: the model must consider every existing thread before raising anything new, and an empty list of new comments is treated as a good outcome, not a failure.

Review depth

--hardness N (-H, default 0) says how hard to look. It is one knob with four settings, and it lives entirely in hardness.py: each level is a Level holding a name, a one-line blurb for the UI, and an addendum that reviewer._build_review_prompt appends to the system prompt.

LevelNameWhat changes
0lightDefault. Only what would change the author's mind or the shape of the code. At most five comments, no style nits.
1standardA full pass over every changed file.
2thoroughAdds a checklist: edge cases, error paths, resource handling, invariants, security, tests, real performance problems.
3exhaustiveAdversarial audit, and the review runs inside a read-only checkout so Claude can read around the diff.

Two properties are worth preserving if you touch this. Depth is orthogonal to the prompt: it says how hard to look, not what to look for, so it is appended on top of a custom prompt file rather than replacing it. And the level never changes the response contract: the same JSON, the same severities, the same approval loop at every setting, so nothing downstream has to know which level ran.

Level 3 is the only one that changes machinery. hardness.explores(level) is true from EXPLORE_FROM up, and reviewer._ask_in_checkout then fetches the PR head, adds a temporary detached worktree, and runs the turn there with Read/Glob/Grep and the credential denylist, the same guardrails as the codebase scan. If the checkout cannot be made, it falls back to reviewing the diff alone rather than failing: the checkout is a convenience, not the review.

Adding a level means adding one entry to LEVELS. The CLI help, the menu, the MCP tool description and clamp() all read the dict, so nothing else needs touching.

Suggested changes

When a fix is small and mechanical, the prompt asks for a ```suggestion block instead of a description of the fix. GitHub renders that as a Commit suggestion button. This is prompt work plus one plumbing change: review comments can now span a range.

  • comments[].start_line is optional in the schema. github.post_pr_review_comment sends start_line/start_side only when it is strictly less than line, because GitHub rejects anything else, so an equal or larger value silently degrades to a single-line comment rather than erroring.
  • reviewer.has_suggestion is the one place that decides whether an item carries a suggestion. The approval loop uses it twice: to mark the item 🔧, and to print the body unwrapped, since textwrap.fill would destroy the indentation GitHub commits.
  • The prompt is explicit that a wrong suggestion is worse than none. A committable-looking bad fix is the failure mode here, so anything uncertain is meant to come back as prose.

Auto mode

praline auto [PR ...] runs the same review as the interactive menu, minus every prompt: auto.run_auto picks PRs, reviews each with reviewer.review_pr, and posts every proposed item straight through reviewer.post_accepted_comments. There is no approval loop.

  1. A PR qualifies if its number was passed as an argument, or if it is not a draft and either has never been reviewed here or has been commented on since it last was.
  2. auto._has_new_activity checks the PR's own updated_at first and only pays for the two comment endpoints when that says no. The fetch exists solely to catch a comment edited in place without the PR being bumped.
  3. Per-user "last reviewed" timestamps live in .praline/auto_state.json (config.load_auto_state / save_auto_state), keyed by PR number, and are only updated after a PR is actually processed.
  4. Before reviewing anything, it refetches full stats per candidate (the list endpoint omits changed_files, same gotcha as get_pr_status) and sums them. Past --max-changed-files it asks a single yes/no confirmation; this is the only prompt auto mode ever shows.
  5. Each review is flattened by reviewer.flatten_review, the same call run_approval_loop makes, and every item is posted unconditionally.
  6. Each PR gets the same tail steps as an interactive review: log, re-request (unless --no-request-review), notify.
  7. A per-PR result (files changed, verdict, comments added, replies left, threads resolved, or the error if the review call failed) is collected and printed as a summary table at the end, and with --slack DMed to you as a round-up grouped by verdict.

This is the one intentional exception to the "no auto-posting" invariant below: a human still approves every comment in the interactive menu, but auto mode exists precisely to skip that step for PRs the user has pointed it at.

What starts a review

select_prs takes on_new_commits, which picks between two predicates. _has_new_comments is the default in both unattended modes: a PR requalifies only when someone has commented. _has_new_activity is the wider one, reached with --review-new-commits, where a push counts too.

The narrow one is the default because nobody asked for the review a push would trigger. An author iterating on a branch would collect a near-identical review per commit, which is noise for them and budget for the user. A comment is someone addressing the review, and that is worth answering.

Both predicates share _commented_since and differ only in how they use pr.updated_at. For the wide one it is a cheap positive: anything that bumps it counts, so the comment endpoints are only paid for when it says no. For the narrow one it is a cheap negative: a PR nothing has touched cannot have been commented on, but a bump alone proves nothing, so the comments are still read. Do not collapse them into one shortcut; the asymmetry is the point.

Monitor mode & budgets

praline monitor is a loop around auto.run_auto, which already picks the PRs worth reviewing. monitor.py adds only what turns a one-shot command into something you can leave running. Keep that division: selection and review logic belongs in auto, not in the loop.

What requalifies a PR is decided in auto.select_prs and shared by both modes: see Auto mode.

The one change monitor mode forced on auto is run_auto(..., confirm_budget=False). The changed-file check is the only interactive moment in that path; unattended, a batch over the limit is skipped and left for a human rather than confirmed by default. run_auto also returns its per-PR results now, which is what the loop reports on.

Each pass empties the comment cache first (github.forget_all_comments). That cache is invalidated on a post and otherwise lives for the process, which is right for a command that exits and wrong for one that looks a hundred times: without clearing it, a comment left between two passes would never be seen and a re-review would be written against a conversation that had moved on. If you add long-running behaviour anywhere else, it needs the same treatment.

The knowledge base is refreshed when PRs merge while watching, via memory.merged_since_last_update. That refresh redraws the module map only when the repo has never had one: the map is the most expensive call PRaline makes and codebase shape moves far more slowly than merged PRs, so redrawing an existing one unattended would spend most of the budget on the part that changed least. A repo only ever watched would otherwise never get a map at all, which is why the first one is exempt.

The page is re-rendered at startup by memory.rerender_html, which reads the documents already on disk and calls nothing. The page and the documents age at different rates: documents change with the repo, the page changes with PRaline's template. Without this, a knowledge base built before a template change keeps producing the old page, and the symptom is a section the code plainly emits being absent from the file the user is looking at.

The budget

A rolling window, enforced before a call, not after. budget.guard() runs at the top of claude_client.ask, so an exhausted allowance means a review never starts, rather than one being abandoned half-posted. auto checks it once per PR and breaks out cleanly, leaving the rest to requalify next run.

Two details that look like details and are not. Cached reads are excluded from the count (budget.billable): a cached system prompt reports tens of thousands of read tokens on a trivial call, at a fraction of the price, so counting them would exhaust a sensible cap in four calls and describe nothing real. And spending is persisted to .praline/budget.json: a monitor that crashes and restarts must not receive a fresh allowance, which is exactly when a cap matters. A stored window that no longer matches the requested one is discarded rather than reinterpreted.

DEFAULT_TOKENS_PER_HOUR is 40,000, and monitor mode is the only mode that gets a budget without being asked, because it is the only one that runs unattended for hours. The number is measured, not guessed: one review of a small PR at depth 0 costs about 12.3k billable tokens, so the default is roughly three an hour. If you change the default, re-measure rather than reasoning about it; --tokens-per-hour takes an explicit 0 to disable.

The active budget is module state (budget.ACTIVE) rather than a config.Run field, because claude_client.ask is called from a dozen places that have no business knowing about budgets. Default is None, and then guard() and record() do nothing at all.

What's new, and PR order

watch.py answers two questions that have nothing to do with the model.

  • What moved since last time. classify tags each open PR NEW (absent from .praline/seen_state.json), UPDATED (its updated_at moved) or SEEN, and run_check prints the digest behind praline check (--no-mark to report without acknowledging) and menu entry 4. Only those two write state (mark_seen); the interactive PR list reads it and never acknowledges, so the icons stay put until you say so. First run has nothing to compare against, so everything is NEW exactly once. is_first_look keys off last_checked, not the PR map, because a repo with no open PRs stores an empty one.
  • What order to review in. order_prs sorts by PR number ascending, except that a PR whose base branch is another open PR's head branch is a stack and comes after its parent. Stacks are walked depth first so they stay contiguous. A cycle or a parent that isn't open falls back to number order; nothing is ever dropped. stack_parents builds the whole parent map in one pass and feeds the ↳ on #5 markers in the list.

Both the CLI picker and auto.select_prs go through this module, so ordering and icons cannot drift apart between the two modes.

Slack notifications

Entirely optional and off unless --slack is passed. cli._load_slack resolves the config once at start-up and exits if it is broken; after that every Slack failure is printed and swallowed, because the review it would report on is already posted.

  • Config. slack.load_config searches $PRALINE_SLACK_CONFIG, then ~/.config/praline/slack.json, then <repo>/.praline/slack.json. The token may come from $SLACK_BOT_TOKEN / $PRALINE_SLACK_BOT_TOKEN instead, so it need not touch a file at all. Nothing here ever writes a token.
  • Who gets the message. notify_review posts to the PR author and the reviewer (cfg.reviewer_login, the GitHub login PRaline authenticated as). Two recipients means conversations.open and a group DM, which is the reason mpim:write is required; one means a direct chat.postMessage to the member ID, which needs no extra scope. Unmapped people are skipped, not fatal.
  • Resolution is cached. A mapping value may be a member ID, an @handle or an email; handle lookups page through users.list, so results are memoized on the config object for the run.
  • The round-up. notify_digest DMs the reviewer a list of everything just reviewed, grouped by verdict, ready-to-approve first. Sent at the end of run_auto, and on quitting an interactive session in which anything was reviewed. notify_digest itself returns early on an empty list, so callers do not need to guard the count.

The review log

.praline/review_log.json is PRaline's memory of its own work: one entry per posted review with the PR number, title, author, head sha, branches, verdict, the overview text and the comment counts, capped at the last 200 by config.append_review_log.

reviewer._format_review_log feeds the newest entries back into later review prompts, always including earlier passes over the same PR and labelling them. That is what makes a re-review build on the last one instead of repeating it, and what lets a stacked PR be read in light of the PR below it, which was reviewed minutes earlier. It is distinct from the knowledge base (lessons from merged PRs) and from seen_state.json (what is new since you last looked).

Huge PRs

GitHub's API refuses to render a PR past roughly 300 changed files as a single diff: get_pr_diff gets a 406 Not Acceptable and raises DiffTooLargeError. reviewer.review_pr catches it and hands off to _review_huge_pr, which warns the user ("huge PR, careful") and recovers in two stages:

  1. Local diff. github.fetch_pr_refs fetches the base branch and the PR's hidden head ref (pull/N/head), then get_pr_diff_local produces the same unified diff with git diff origin/<base>...<head_sha>, no size limit. The PR's true size (from git diff --stat, since PRInfo counts are 0 off the list endpoint) is printed before going further.
  2. Explore mode. Interactively, the user is asked; in auto mode it is enabled directly, since auto mode has no prompts. github.add_pr_worktree checks the PR head out detached in a temporary worktree, the diff is written to PRALINE_DIFF.patch at its root, and the review call runs with readable set to the worktree and the scan's longer timeout. prompts.HUGE_PR_EXPLORE_ADDENDUM tells the model to page through the diff in slices and use the checkout for context. The worktree is force-removed in a finally. Declining explore mode inlines the full local diff instead, which works whenever it fits the model's context.

Either way the response contract is untouched: the same JSON schema comes back, and the approval loop or auto-posting proceeds as for any other PR. The feature holds the four holy principles (all six of them): MINIMALISM, nothing changes before a 406 and the tooling is the scan's, reused; CLEANLINESS, one prompt, one schema, only the diff's transport changes; MODULARITY, git plumbing in github.py, flow in reviewer.py, instructions in prompts.py; EFFICIENCY, the diff stat steers the model's reading budget toward the files that matter; SAFETY, the worktree is temporary and detached, so the user's checkout, index and branches are never touched; SECURITY, explore mode runs read-only with the credential denylist, no shell and no network, exactly like the codebase scan.

The knowledge base

Menu option 2 calls memory.update_knowledge, which returns the (markdown, html) paths so the CLI can print them. It writes into .praline/ inside the reviewed repo, not inside PRaline:

FileWritten byContents
repo.mdscan_codebase or build_repo_knowledgeArchitecture, module reference, conventions, invariants, pain points.
pr_history.mdbuild_pr_historyLessons from merged PRs in the chosen window, every claim cited as (#123) and linkified to the PR page.
knowledge.mdrender_knowledge_mdThe two halves stitched together under one title, for reading and grepping.
knowledge.htmlrender_knowledge_htmlThe same document through templates/knowledge.html, standalone and publishable.
knowledge_state.jsonupdate_knowledgeBuild timestamp, used to detect newly merged PRs.
artifact_url.txtyou, optionallyPointer to a published copy of the HTML.

repo.md is built two different ways, and which one runs is decided by whether it already has content:

  • First build: scan_codebase. A tool-enabled Claude turn runs with cwd set to the repo and reads it directly (Read, Glob, Grep). It follows entry points and imports rather than guessing from a file list, so this is the pass that produces the module reference and the invariants. Slow, and it runs once. Timeout is EXPLORE_TIMEOUT_S, not the default.
  • Every later build: build_repo_knowledge. No tools. It looks up the default branch via the GitHub API, runs github.fetch_remote_branch, and reads git ls-tree / git log off the resulting origin/<branch> ref instead of HEAD, so a stale local checkout cannot skew it. git fetch only updates refs/remotes/*, never the working tree, the index, or local branches. If your checkout was behind, a one-line note reports by how many commits; if the fetch failed outright, it says why and falls back to reading HEAD.

Fetching goes through one function, github.fetch_refs, and it tries twice. First git fetch origin, using whatever credentials the remote is configured for. If that fails it retries over HTTPS with the GitHub token PRaline already requires. That second attempt is not a nicety: the remote is usually ssh, and an ssh agent is the one thing PRaline cannot assume is there: a tmux session outliving the agent that started it, or a machine with no key loaded, used to leave every fetch dead behind a bare "could not fetch origin/main" while a perfectly good token sat in the environment.

The token reaches git through the environment and is referenced by name in the credential helper, so its value never appears in argv where ps would show it to every user on the machine, and never lands in .git/config. An empty credential.helper= is passed first to reset the helper list, so a system helper configured elsewhere cannot answer ahead of it with the wrong account. Both attempts use the same explicit refspecs, so what ends up in refs/remotes/ does not depend on which one won. Only a failure of both raises FetchError, carrying the first line of git's own message, run through redact_url, since that message may quote the remote.

Both markdown halves are fed into the review prompt on every review, which is the whole point: reviews get repo-specific over time.

Updates are edit passes, not rewrites. The previous document is included in the prompt with instructions to merge, and _guard_against_erasure refuses any result shorter than half the previous version, keeping the old file instead. A .bak copy is written before each save. If you touch this flow, preserve both protections; silently losing accumulated knowledge is the worst failure mode this module has.

Staleness is checked on the way into a review. memory.last_updated_at reads the build timestamp (falling back to repo.md's mtime for knowledge bases built before that file existed), merged_since_last_update asks GitHub what has merged since, and cli._offer_update_for_new_prs lists those PRs and offers to refresh. A failure here prints a warning and gets out of the way; it must never block the review the user actually asked for.

The module map

knowledge.html ends with a "Repo at a glance" section: stat tiles over a node-link diagram of the modules. Three pieces, deliberately separated.

  1. Getting it. graph.build runs one read-only Claude turn against the checkout with prompts.ARCH_GRAPH_PROMPT, which asks for {nodes, edges} and nothing else.
  2. Trusting none of it. graph.normalize drops what cannot be drawn: unknown node kinds fall back to module, duplicate ids, self-loops, repeated pairs, and any edge naming a node that was never declared. It never raises. A diagram missing two boxes is worth showing; a knowledge base that failed to build because the map came back malformed is not. memory.build_graph takes the same line one level up: on any failure it keeps the previous map and carries on.
  3. Drawing it. render.py emits the section and embeds the graph in a <script type="application/json"> block; the template's JavaScript lays it out. No CDN, no build step, so the page stays a single self-contained file you can publish as an artifact.

The layout is a layered DAG, not a force simulation, so the same graph always draws the same picture. Cycles are broken before layering, with a depth-first walk that marks the edges closing a cycle; those are excluded from the layer assignment and rendered dashed. Skipping that step is not cosmetic: relaxing longest-path across a cycle pushes its nodes a layer deeper on every pass, and four mutually-importing modules end up strung across a canvas wide enough for thirteen.

The stat tiles come from github.repo_counts, never from the model. That function lives in github.py because every git call does; graph.stats only attaches labels. A number a model wrote down is a number nobody can check. Tiles that come back zero are dropped, so a shallow clone renders fewer tiles instead of a row of noughts.

Node kind is carried by shape and outline (a ring and ▸ for an entry point, dashed for external), not by color alone, and a table view sits under the diagram. If you restyle it, keep both: the theme's pastels are decorative and do not survive a colorblind-separation check as a categorical encoding.

The MCP server

praline-mcp exposes the same operations as MCP tools, so Claude can drive PRaline directly. It is an optional extra (uv sync --extra mcp); nothing else in the package imports mcp_server, so the dependency stays optional in fact and not only in the metadata.

Each tool resolves its repo the same way, through _resolve: the explicit repo_dir argument, else PRALINE_REPO_DIR, else the process's working directory, which is the repo you are in when Claude Code launches the server.

Two constraints shape the module, and both are easy to break by accident:

  • Drafting and posting are separate tools. review_pr_draft never writes to GitHub; it stashes the result in _DRAFTS and returns the numbered draft. post_review posts the indices the user approved, out of that stash rather than by re-running the review, which would come back subtly different. A model that reads "review this PR" as "review and post it" therefore cannot post anything.
  • stdout belongs to the protocol. The rest of PRaline narrates as it works, and on a stdio transport one stray print corrupts the JSON-RPC stream. Every tool body runs inside _captured(), and _reply() hands the captured text back as the result's log so the progress is not simply lost. If you add a tool, wrap it the same way.

The Claude backend

claude_client.ask() is the only place PRaline talks to a model. It runs:

claude -p --output-format json --model <model> \
       --allowedTools <scoped-read-rules> --disallowedTools <secret-globs> \
       --setting-sources "" --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
       --name <session name> --system-prompt-file <temp file>

Details worth knowing before you modify it:

  • Nothing that grows with the repo goes into argv. Linux caps a single argv entry at 128 KiB (MAX_ARG_STRLEN), whatever getconf ARG_MAX reports, and both of the big strings pass that on a real repo. The user message (typically a large diff) is piped through stdin. The system prompt carries the knowledge base, the PR history lessons and the review log, so it grows with the repo; it goes through a 0600 temp file written by _write_system_prompt and deleted in a finally. Passing either as an argument fails with [Errno 7] Argument list too long, and it fails later the smaller your knowledge base is, which is why it first showed up in monitor mode. If you add another argument, keep it bounded.
  • Every turn is named, and the name is required. session_name is what the session is called in /resume, the prompt box and the terminal title, instead of a hex id. Callers pass the tail only; ask() prepends SESSION_PREFIX, so one grep finds every session PRaline started and no call site can spell the prefix differently. Reviews get PRaline: owner/repo PR #12 (review 2) from reviewer._session_name(), which counts earlier passes over that PR in the review log; the knowledge base, PR history, codebase scan and module map name themselves. Build a name from what PRaline already knows, never from PR text: a name goes into argv, where the diff deliberately does not, and a PR title is written by whoever opened the PR.
  • There are exactly two call shapes, and the confinement is structural. ask() takes readable, one directory, and there is no tools= parameter any more. Omit it, as reviews and update passes do, and the turn gets no tools. Pass it, as the codebase scan, the module map and explore mode do, and it gets Read, Glob and Grep scoped to that directory by _read_only_rules, plus the secret denylist and _ISOLATION_ARGS, which ask() applies itself so no call site can omit them. A bare Read allows reading any path on the machine, which is why the old signature was removed rather than documented. Scoping every read tool matters: Grep takes its own path and returns the matching lines of a file it was never allowed to open. Nothing else should need tools; if you think it does, see Secrets and data flow first.
  • No settings, no MCP servers. Explore mode runs with the working directory set to a checkout of the PR under review, so a .claude/settings.json there would install hooks and a .mcp.json would launch a server command: arbitrary code execution triggered by opening a pull request. _ISOLATION_ARGS closes both, and keeps a review away from the user's own MCP servers. Never restore a settings source to widen a permission; scope the rule instead.
  • Enforcement is the permission layer, not the model. Verify by asking a turn to read outside its directory and finding the attempt in the reply's permission_denials. A model declining is not evidence of anything.
  • Every failure mode (timeout, non-zero exit, empty stdout, bad JSON, error payload, empty result) raises a RuntimeError carrying truncated stderr/stdout. Keep that style: errors here surface directly in the CLI and are the user's only debugging signal.
  • The default timeout is 600 seconds; big diffs on big models are slow. The scan gets 2,400.
  • The subprocess runs with env=_clean_env(): the current environment minus GITHUB_TOKEN, GH_TOKEN, SLACK_BOT_TOKEN and PRALINE_SLACK_BOT_TOKEN. The turn has no use for them, and a process that never receives a secret cannot leak it, whether through a tool, a hook, an MCP server, or a diff written to talk the model into echoing its environment. Add any new credential variable to that tuple the moment you introduce it.

Secrets and data flow

The knowledge base is the one place where reading and publishing meet. Whatever lands in repo.md is fed back into every later review prompt, can be posted to GitHub as part of a comment, and may be published as an artifact. So the scan is constrained on two levels, and both matter:

  • The CLI blocks credential files. The secret denylist in claude_client.py is a list of Read(<glob>) deny rules (.env*, *.pem, *.key, id_rsa*, .netrc, secrets.*, credentials.* and more). They match credential files on purpose: a blanket *token* would hide tokenizer.py, and a scan that quietly skips real code is its own failure. Deny rules override the allowlist and apply to Grep as well as Read, so a blocked file cannot be laundered through a content search. This is enforced by the harness, not by the model's judgement.
  • The prompt handles what globs cannot. A key hardcoded in ordinary source is not matched by any denylist, so INIT_CODEBASE_PROMPT requires recording that such a secret exists and where, never its value.

If you extend the scan, add the glob before adding the capability. Other things that hold:

  • The GitHub token is read from the environment in github._token per request and only ever sent as an Authorization header to api.github.com. It is never logged, never written to .praline/, never interpolated into a message, and never put in a URL. Error paths raise on the response, which does not carry request headers. Same rules for the Slack token, which only reaches slack.com.
  • A token can arrive from somewhere PRaline does not control: a git remote of the form https://ghp_xxx@github.com/owner/repo. github.redact_url strips the userinfo part before that string can reach an error message. If you print a remote URL anywhere, print it through that function.
  • Credentials are stripped from the review subprocess environment; see The Claude backend. The list in claude_client._CREDENTIAL_ENV_VARS includes FETCH_TOKEN_ENV, the variable github.fetch_refs passes the token to git in. It is defined next to the strip list and imported from there, because naming it in two places is how it ends up stripped in only one.
  • cli._warn_if_praline_dir_committable runs git check-ignore at start-up and warns when .praline/ is not ignored in the target repo. That directory holds the knowledge base, the review log and possibly a Slack config.
  • System prompts are passed on argv, so they are visible to other users on the same machine via the process list. That covers the knowledge base and the diff-free part of a review, which is repo-internal but not secret. Do not move anything sensitive into that argument.
  • .praline/ lives in the reviewed repo and describes it in detail. PRaline warns when it is not ignored but cannot fix it, since it never commits.
  • A PR diff or description is untrusted input that ends up in a prompt, in the same prompt as the knowledge base, the PR-history lessons and the review log. The interactive approval loop is what stands between that and a posted comment. auto and monitor remove it deliberately, so on those paths a PR crafted to make the model quote its own context can put internal notes into a public comment. That is the residual risk of unattended review and the reason both are opt-in.

Invariants

These are promises the README makes to users. A PR that breaks one will be rejected regardless of how useful the feature is.

  • No write-to-code endpoints. Every function in github.py maps to a read endpoint, a comment endpoint, or the review-request endpoint. When adding a GitHub call, keep it inside the ALLOWED list documented at the top of that file, and note what is not there: approving, merging, and opening or editing issues.
  • No auto-posting outside praline auto. The interactive menu's comments flow through run_approval_loop and the final confirmation; nothing in cli._do_review may skip them. auto.run_auto is the one deliberate exception, gated behind its own subcommand.
  • No API key. The model backend stays the claude CLI. Do not introduce anthropic SDK calls or key handling.
  • No tools beyond read-only, and never unscoped. A tool-enabled call names its directory through readable and gets nothing else. A write or shell tool would break the "never touches code" promise outright, and an unscoped read tool hands a prompt-injected PR the whole filesystem.
  • Credentials never reach a subprocess or a log line. Tokens are read from the environment at the point of use, sent only to their own API host, and stripped from the claude child environment. Nothing writes one to disk, and no error message may interpolate one.
  • Slack only ever sends, and only to people. The four methods called are auth.test, a user lookup, conversations.open and chat.postMessage to a DM or group DM. No channels, no history reads, and a Slack failure never rolls back or blocks a review.
  • Nothing posts from the MCP server without a second call. review_pr_draft drafts and returns; post_review publishes what the user picked. Any new tool that writes to GitHub stays separate from the tool that produces the content, for the same reason the interactive loop exists.
  • Every review is grounded in the accumulated knowledge. reviewer._build_review_prompt is the single place that appends repo.md, pr_history.md and the review log, and every review path goes through it, including huge-PR explore mode, depth 3, and the MCP tool. A new path that assembles a prompt some other way silently loses the repo-specific half of the review. The module map is deliberately not in there; it is for the HTML page.
  • The rendered knowledge base never carries live markup. render._NoRawHtml deregisters Markdown's html_block preprocessor and html inline pattern, so raw HTML in repo.md or pr_history.md becomes visible text. Both files are model output over content that includes attacker-controlled PR titles and bodies, and the page they produce is opened over file:// and published as an artifact. Do not "fix" this by escaping the HTML stash instead: fenced_code shares that stash, and every code block in the document would be double-escaped.
  • Anything untrusted printed to the terminal goes through term.plain. PR titles, author names and comment bodies are written by other people, and GitHub does not strip control bytes on the way out. A title carrying \r\x1b[2K erases its own line and reprints whatever the author chose, which in a tool driven by [a]ccept/[r]eject and [Y/n] prompts is a way to forge what the user thinks they are approving. Colour is applied by _c around the sanitized text, never taken from the text. If you add a print of anything that came from GitHub or from the model, wrap it.
  • Anything quoted into a Slack message goes through slack._escape. Slack reads <...> as markup, so an unescaped PR title can forge a link or ping the room. The <url|label> syntax this module builds is applied after escaping, never to it.
  • Unattended modes stay capped and non-draft. Monitor mode gets a token budget whether or not one was asked for, and drafts never qualify for an automatic review. Both are what makes leaving it running defensible.
  • The review reply is JSON with a fixed schema. The schema lives in prompts.DEFAULT_REVIEW_PROMPT and its consumers are run_approval_loop and post_accepted_comments. Change all of them together, and remember users can supply custom prompt files that must still produce the same schema.

Where changes go

You want to...Touch
Add a menu actioncli._main_menu plus a _do_* helper next to the existing ones; take a config.Run rather than loose arguments.
Add a global setting or flagcli._add_common_args for the flag, a field on config.Run for the value. Both modes then see it without any signature changes.
Change how reviews behaveprompts.py first. Most behavior (tone, priorities, reply-before-comment) is prompt-defined, not code-defined.
Add a GitHub operationgithub.py, respecting the ALLOWED list. Return plain data (dataclasses or dicts); rendering belongs to callers.
Change what a review seesreviewer.review_pr and _format_existing_conversation for the user message, _build_review_prompt for the system side.
Change the approval UXreviewer.run_approval_loop, _display_comment, _prompt_action.
Change knowledge-base contentprompts.INIT_CODEBASE_PROMPT (first build), INIT_REPO_PROMPT (updates) or INIT_PR_HISTORY_PROMPT; the surrounding pipeline is memory.py.
Change how the knowledge base readsprompts.KB_STYLE. It is appended to all three knowledge-base prompts, so the rules stay in one place.
Restyle the knowledge-base pagepraline/templates/knowledge.html for the shell and the map's JavaScript, render.py for the fragments Python builds.
Add or change a review depthhardness.LEVELS, and nothing else. The CLI help, the menu, the MCP tool description and clamp() all read the dict.
Change what the module map showsprompts.ARCH_GRAPH_PROMPT for what Claude is asked, graph.normalize for what survives, the template's script for how it draws.
Add an MCP toolmcp_server.py: a @server.tool() function wrapping its body in _captured() and returning _reply(...). Keep anything that writes to GitHub behind its own explicit tool.
Add a stored file under .praline/config.py for the path and load/save helpers, then use them from wherever needs it.
Change auto-mode selection or postingauto.select_prs for which PRs qualify, auto.run_auto for the review/post/summary loop. Monitor mode inherits both.
Change what monitor mode does per passmonitor.run_monitor for the loop, monitor._report for what a pass prints. Anything about which PRs qualify belongs in auto.
Change the review order or the new/updated iconswatch.order_prs and watch.classify. Both modes read them, so a change lands everywhere at once.
Change a Slack messageslack.review_message (per PR) or slack.digest_message (the round-up). Delivery is notify_review / notify_digest; keep failures non-fatal.
Add or change a verdictverdict.STATUS_LABEL and verdict.status_key, plus the status guide in prompts.DEFAULT_REVIEW_PROMPT. The three renderers (CLI, auto summary, Slack) all read the same helpers.
Record something about past reviewsreviewer.log_review for the entry, _format_review_log for what the prompt sees.

Gotchas

  • A model reply that should be JSON goes through claude_client.extract_json. It tolerates a fenced block or prose either side of the object, and it is shared by the review parser and the module-map parser. Do not write a third one.
  • The module map is laid out from a cycle-free graph. graph.normalize removes self-loops but not cycles; the template breaks those at layout time. If you rewrite the layering, keep the depth-first cycle break or a mutually-importing package will draw itself across a canvas ten times too wide.
  • A suggestion block must not be re-wrapped. Its indentation is what GitHub commits, so _display_comment prints those bodies verbatim. Any new renderer of comment bodies has to do the same.
  • Resolving review threads is GraphQL-only. GitHub's REST API cannot mark a conversation resolved, so github.py carries two small GraphQL calls (find_review_thread_id, resolve_review_thread). Everything else is REST.
  • reply_to_id must be the thread's root comment. Replying to a reply detaches the comment from its thread. The prompt instructs the model accordingly; keep that instruction if you rewrite it.
  • get_merged_prs paginates by updated, not merged_at. The early-exit condition is deliberately conservative (a whole page must be older than the cutoff) and there is a hard cap of 10 pages. Easy to break if you "simplify" it.
  • Claude sometimes wraps JSON in code fences. reviewer._extract_json tries the raw text, then a fenced block, then the outermost {...}, and keeps the first that parses. Removing this "redundant" code will produce intermittent parse failures.
  • The same goes for prose around a document. memory._strip_preamble drops anything before the first markdown heading, because a scan that used tools sometimes opens with a sentence about what it read. The prompts forbid it; the strip is the backstop.
  • The HTML template uses safe_substitute, not substitute. It contains literal $ characters in its prose, which substitute would reject as bad placeholders. Keep the placeholders ($title, $repo, $toc, $body) distinctive if you edit it.
  • The knowledge files live in the reviewed repo. Consider suggesting .praline/ for that repo's .gitignore; PRaline itself never commits anything, so it cannot do this for the user.
  • get_repo_structure / get_recent_commits take an explicit ref. They read git ls-tree / git log <ref>, not the working tree or HEAD. Callers outside build_repo_knowledge must still pass a ref (e.g. "HEAD") if they want the local checkout.
  • Subparser defaults must be SUPPRESS. _add_common_args is called on the top-level parser and on every subparser so --dir works on either side of the subcommand. argparse applies a subparser's defaults after the top-level parse, so a real default on the subparser silently overwrites what the user passed before the subcommand, which is why _add_common_args(..., subcommand=True) swaps every default for argparse.SUPPRESS. Add a new shared flag through that function, never directly to a subparser.
  • Comment lists are cached per run. github._COMMENT_CACHE makes the three readers of a PR's comments (status line, prompt conversation, reply lookup) cost one fetch each instead of three. Every posting function calls forget_pr_comments; if you add another write path, call it too or the next read will be stale.
  • Line comments take the head sha from the caller. post_pr_review_comment no longer refetches the PR per comment. Pass pr.head_sha, which also pins the comment to the commit that was actually reviewed.
  • You cannot request a review from the PR author. GitHub answers 422, so rerequest_review skips your own PRs. Anything else it fails on is printed and ignored.
  • PRInfo counts are only populated on single-PR fetches. The list endpoint omits additions/deletions/changed_files, so those default to 0 until get_pr is called; that is why get_pr_status and auto.run_auto both refetch. draft and updated_at, by contrast, are present on the list endpoint already.