Reviewing What a Memory-Fed Autofix Taught¶
Agentic autofix stores its fix pattern before any merge decision, so the merge review decides whether the run teaches the repository.
A repair agent that writes to a shared memory store makes two changes per run, and the pull request shows one of them. Agentic autofix reads Copilot Memory for context on a security alert. "When it creates a fix, it stores the fix pattern as a memory for future use" (GitHub Changelog, 2026-09-25). Three other surfaces read what it wrote, because "Copilot Memory is currently used by Copilot cloud agent, Copilot code review, Copilot CLI, and agentic autofix" (GitHub Docs).
When the write is worth reviewing¶
Two conditions have to hold. Miss either and auditing the write buys nothing.
- Whoever assigned the alert has memory on. Facts are created only "in response to actions by users with write access to the repository who have Copilot Memory enabled", and the switch "is enabled per user, not per repository" (GitHub Docs). Two runs on one alert differ in whether they teach the repository.
- The fix merged. Facts "can also be captured from pull requests that were closed without merging". A pattern from an abandoned branch then faces a further test: "the validation step ensures that Copilot's behavior is unaffected unless the current codebase still substantiates the information" (GitHub Docs).
A third fact sets how much the green CodeQL check tells you. Autofix "can't confirm that a fix resolves alerts generated by custom queries or the security-extended query suite" (GitHub Docs). For those alerts the check is not evidence that the fix works, so the review carries the whole test. Classification before repair is the earlier gate there.
What the validation step confirms¶
Every repository-level fact carries citations into the code that supports it, which is the mechanism Copilot Memory uses to keep itself current. Copilot checks those citations "against the current branch to confirm the information is still accurate. Only validated facts are used" (GitHub Docs). That test asks whether the fact still describes the code. It never asks whether the fix was correct.
Merging inverts which case is safe. A rejected pattern loses its substantiation when the branch closes. An approved one gains it, because the code now exhibits the pattern the memory describes, so validation passes by construction. Then use keeps the entry alive. "The 28-day timer may reset whenever Copilot successfully validates and uses an entry", while an unused one is "automatically deleted after 28 days" (GitHub Docs).
An approved fix can still be wrong in ways CodeQL does not see. Autofix "may suggest fixes that only partially address the security vulnerability or only partially preserve intended code functionality" (GitHub Docs). GitHub's mitigation against that is scoped to the diff: autofix "presents all suggestions as proposed code changes that require explicit developer review and acceptance before being applied". The same application card's mitigation list does not mention memory anywhere.
So add one question before you merge. Which convention does this fix assert, and do you want Copilot code review enforcing it next month? Where the answer is yes, write it into the repository's custom instructions, because Copilot follows them when it generates a fix (GitHub Docs). Where it is no, read the repository's stored facts.
Why it works¶
Retrieval gives a repair agent precedent the alert text does not carry. Without it, EvoRepair reports, models "may repeatedly make similar mistakes during iterative repair and underutilize valuable repair knowledge from historical vulnerabilities" (arXiv 2605.30105v2, from the abstract; arXiv offers no HTML rendering of this paper). The same channel carries bad precedent. Persistent memory "makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes" (arXiv 2608.10502v1).
Catching the pattern at merge time therefore beats removing it later. The only remediation GitHub documents is manual: "Repository owners can review and manually delete the repository-level facts stored for their repository" (GitHub Docs). Deletion alone recovers little. On a 150-case controlled benchmark, a delete-the-retrieved-memories baseline recovered 34.0% of cases against 85.3% for dependency-guided rollback. It removed 88.1% of the diagnosed faulty memories rather than all of them (arXiv 2608.10502v1, Table 2). That is a research harness and not Copilot, so read it as a bound on the mechanism.
When this backfires¶
- No contributor has memory enabled. Nothing is written, and the extra review question is pure overhead.
- The repository runs autofix a few times a quarter. The 28-day expiry prunes an unused entry before any pattern accumulates.
- The reviewer is not a repository owner, so the follow-through above is unavailable to them.
- The convention is already written down. Instructions take effect at generation time, which costs less than policing what a run inferred.
- Reviewer attention is the binding constraint, which is the strongest case against this technique. Across 33,596 agent-authored pull requests in GitHub repositories with at least 100 stars, 61.38% "receive no recorded review activity" (Duma et al., 2026). The paper adds that "the absence of recorded review activity does not imply the absence of human oversight". Where reviewers leave no trace on the diff, asking them to audit the store as well adds load.
Example¶
GitHub documents that autofix can fix an alert only partly. That limitation looks like this in a review queue. A path-traversal alert is assigned to Copilot. The agent explores the repository, adds a normalization helper in one route handler, re-runs CodeQL, and opens a pull request. CodeQL is satisfied, so the reviewer sees a green check and a small diff, and merges it.
The helper covers that handler and not the other four that read user-supplied paths. The stored fact says the repository normalizes user paths through this helper, and it cites the merged lines, which do exactly that. Every later validation passes. "Copilot code review uses repository-level facts only" (GitHub Docs), and this fact is one of them. So the next pull request touching a sibling handler gets reviewed against a convention that holds in one file.
The review that catches this asks one question before merging: does this fix state a repository-wide convention, and is it true repository-wide? Here it is not. The fix ships with an issue for the remaining four handlers. The reviewer reads the stored facts instead of leaving the claim to be validated by the line it just created.
Key Takeaways¶
- The write precedes the review decision, so closing the pull request does not delete the stored pattern, though the pattern then loses its substantiation (GitHub Docs).
- Citation validation confirms a fact still matches the code, never that the fix was correct. A merged partial fix therefore substantiates its own memory and renews its 28-day clock on every use.
- Whether a run teaches the repository depends on the person who assigned the alert, because Copilot Memory is enabled per user rather than per repository.
- Autofix's documented human-in-the-loop mitigation covers the suggested code change, and the application card's mitigation list does not mention memory.
- Skip this review where nobody has memory enabled, where the reviewer cannot delete facts, or where the diff itself is going unread.
Related¶
- Copilot Memory and Cross-Agent Persistence — How the store works: entry shape, scopes, citation verification, and the controls you reach for once this review finds something.
- Learned Review Rules — The opposite direction of travel, where accepted and rejected review feedback becomes the persisted rule.
- Review-Feedback-to-Rule Loop — What to do with a convention you decide is real: promote it to a mechanical check instead of leaving it as an inferred fact.
- Classification Before Repair in an Analyzer Backlog — The gate that runs before this one, for a queue where deciding the finding is real is the expensive judgment.
- Agent Memory Patterns — Scopes, what to persist, and why non-obvious corrections are the highest-value entries.