HOW WE PROVE
THE SCORING WORKS.
The Antares scoring engine is closed source — not because we have something to hide, but because publishing the code would also publish every loophole a rug-puller needs to evade detection. That puts us in the same boat as the anti-fraud teams at banks, ad networks, and exchanges: credibility has to come from outcomes, not from the codebase.
This page documents what we test against, how often, and what an independent reviewer can verify without ever seeing a line of source.
The backtest corpus is a library of real Solana token scans, captured by hitting the exact same /api/scan endpoint your extension hits. Each entry records the verdict the engine returned, the score it produced, every flag it raised, and the supporting on-chain state at capture time — holder count, liquidity, LP-burn status. Nothing is synthesised. Every row is a real Solana mint, captured on the day it was scanned, so a row shows how the engine behaved then, not necessarily today.
The corpus leans toward dead and rugged tokens: of the 504 captured entries, 233 are labelled RUG or DANGER, 229 CAUTION and 42 SAFE. Labels come from market data (market cap, liquidity, age, price change), not from an investigation, so a label is a hint, not a verdict. The corpus exists for regression: a change to the engine that flips a stored verdict outside its tolerated band is caught before release. It is not a measure of accuracy, and we do not publish one.
Every Monday an automated job re-scans a rotating sample of 100 corpus tokens against the live /api/scan endpoint (the whole corpus in about five weeks) and compares the new verdict and score against the captured snapshot. It reports three things:
A verdict moved
A re-scanned token comes back with a different verdict than its snapshot. Listed in the report; when more than 5% of the sample moved, a tracking issue is opened and refreshed on later runs.
The check could not measure
More than 20% of the scans failed, or nothing could be compared. The job fails visibly instead of passing green having measured nothing.
Score drift
Same verdict, but the score moved by more than 150 points. Listed in the report so a human can decide whether the label needs an update or the engine needs a tune.
The job hits production — the same endpoint your extension uses. No mocks, no synthetic data, no shortcuts. Requests are spaced out so the check does not compete with real user traffic. The corpus itself is read-only; the job verifies, it never mutates.
Closed source doesn't mean unauditable. The audit happens through behaviour, not through code review. Here is exactly what an independent reviewer can confirm today:
The endpoint is real.
Every public /api/scan call is the same code path the corpus drift check hits. You can compare the JSON your extension receives against a token's known on-chain state and the result is reproducible.
The verdicts have a logic.
Re-scan any token currently in the corpus and the score should land in the same band as the snapshot unless the on-chain state, or our engine, genuinely changed (engine changes are listed in the changelog). When the band shifts for no chain reason, that's a story we owe you.
Flags are explicit, not vibes.
Every flag exposed in the overlay is a named string with a severity. A paying user opens Critical Flags and sees exactly which signal moved the score — not "score 105, trust us".
The check can fail.
A run that finds drift opens a tracking issue, and a run that could not measure fails visibly — not a green tick that means nothing.
For security researchers, integration partners, and exchange compliance teams who need a deeper look than behaviour can give: we hand out scoped read-only access to the scoring repository under a one-page NDA. The review covers the engine, the cross-validation layer, and the Safe Gate logic — everything that turns inputs into a verdict.
Email with subject [AUDIT] and a short description of who you are and what you want to verify. Replies within 48 hours.
Antares is a heuristic risk-screening tool, not an oracle. A SAFE verdict means none of the detection layers found a critical signal at scan time — it is not a guarantee the token won't rug. False negatives happen, especially on novel attack patterns. The whole point of the corpus drift check is to catch them as fast as the corpus reflects the new attack — not to pretend they never occur. Read the terms for the full disclaimer.
The description above is exactly what the test runner does — not a marketing reframe of it. If the corpus grows, the same loop runs against the new tokens. If a labelled rug starts evading the engine, the test fails before the new build ships. If we are wrong, the test fails, and the truth shows up in the same place this page does.