Methodology & limits
Useful enough to help you choose what to review. Not reliable enough to decide a case by itself. That is the standard this product is built around.
An AI code detector can help you sort a large set of files and find the ones that deserve a closer read. It should not be treated as proof of misconduct. A high score means "review this first." It does not mean "this person cheated."
Short programs, starter templates, common libraries, and strict style rules can make human code look generic. Rewrites and mixed authorship can make AI-assisted code look more human. The same submission can look different depending on language, assignment, and prompt design.
A 2024 code-specific study found that existing AIGC detectors performed poorly when separating human-written and AI-generated code.
The report is designed to make the score inspectable. Instead of one number and a shrug, it should show the file-level result, the snippets that contributed most, and enough context for a human to decide what to ask next.
We still do not publish a single accuracy percentage. What we can publish is where our evaluation stands today: a named test set, a date, a method summary, and the failure modes we already know about.
Test set
Our most recent evaluation ran on CoDET-M4, a public code-detection benchmark, plus our own held-out test set — 1,922 code samples in total, human-written and AI-generated, across six languages: Python, Java, C++, JavaScript, HTML, and Lua.
Evaluation date
October 8, 2026. We re-run the evaluation when the model, the scoring threshold, or the input pipeline changes, and this section is updated when a new round completes.
Method
Every sample went through the same input normalization our production scanner uses, and every sample was scored at the same production threshold. That threshold is calibrated so that human-written code in the evaluation set is not flagged. The trade-off is deliberate: some AI-generated code will pass, because a tool that flags honest work causes more harm than a tool that asks for a second look.
Known failure modes
This is a snapshot, not a certificate. The next evaluation round will update this section — including human-side validation for HTML, JavaScript, and Lua.
We will not claim a fixed accuracy percentage without evidence. We will not say a report proves cheating. We will not tell students how to make AI-generated code pass screening. Those claims make the category less trustworthy and make real review harder.
Because one number hides the part that matters: which kinds of code fail, on which languages, and under which edits. Until we can publish that honestly, we will not publish a number.
No. In our most recent evaluation (October 8, 2026), Python, Java, and C++ were tested with both human-written and AI-generated samples. HTML, JavaScript, and Lua had AI-generated samples but no human-written ones, so results in those languages are extrapolated and carry more uncertainty.
A false positive is human-written code that gets flagged. Short assignments, starter templates, and strict formatting can all raise that risk.
A false negative is AI-assisted code that does not get flagged. Rewrites, mixed authorship, and heavy editing can all reduce the signal.
Not necessarily. If you treat the output as a screening signal and keep a human review step, it can still save time and focus attention.
Students can use a free check to understand what might get flagged and to prepare an explanation of their work. We do not provide advice for hiding AI use.