VeriTrust does not publish an accuracy, precision, recall, or F1 claim until a controlled, reproducible benchmark is completed. Runtime risk scores are investigation signals, not measured product accuracy.
No VeriTrust benchmark is currently published.
The repository does not contain a controlled labeled evaluation, confusion matrix, or verified accuracy, precision, recall, or F1 result. Accordingly, the site makes no numerical performance claim.
Current scan output is an AI-assisted triage signal. A provider model's confidence value is not a VeriTrust accuracy measurement and is not legal, forensic, cybersecurity, or final proof.
Published evidence will appear only after an evaluated model card has been approved.
Published model cards
No metric is displayed until its supporting model card is published.
MailGuard email-threat evaluation
The current message paths are VeriTrust MailGuard, a hosted text classifier, and VeriTrust Cortex, a hosted instruction model. The displayed phishing score combines model output with deterministic indicators, so it is not a raw model benchmark.
- A valid evaluation would report email, SMS, URL, and short-message subsets separately.
- It would assess model-only predictions separately from the combined rule score.
- Any benchmark must state that a low-risk result does not prove a message is safe.
Swift URL-intelligence evaluation
The available link path uses a server-configured VeriTrust Swift classifier plus deterministic URL-string indicators. If hosted inference fails, the route returns a disclosed local URL-pattern fallback. It does not fetch pages, resolve redirects, query WHOIS, or use live reputation feeds.
- No VeriTrust link-classification benchmark is currently published.
- Classifier output and deterministic URL indicators would need separate evaluation.
- Fallback frequency and URL-pattern-only outcomes must be measured separately from hosted-classifier results.
Metrics required before a performance claim
- Accuracy, precision, recall, and F1 score on documented labeled test sets.
- False positive and false negative rates by scan type.
- Confusion matrices for each evaluated model and dataset split.
- Latency and fallback frequency under production-like conditions.
Dataset notes
Any future benchmark would need to identify its dataset source, labeling process, sampling approach, train/test separation, known bias, exclusions, preprocessing, and evaluation date. Without those details, a headline accuracy figure would not be reliable.
Limitations
- False positives can incorrectly flag benign content as suspicious.
- False negatives can miss risky content, especially when attackers change wording, domains, or visual generation methods.
- Model fallback can change score distributions because a different model handled the request.
- High-impact cases should include manual review and trusted source verification.