Integrity Scan methodology · ruleset 2.0

A local judgment you can recompute yourself.

Integrity Scan uses a deterministic, non-additive review-priority model. The same supported evidence produces the same result; the number is not a probability of fraud and does not govern Deep Review.

0–34Low review priority

No strong supported file signal. Source verification can still be required.

35–64Moderate review priority

Reconcile the observed signals with the issuer or expected source workflow.

65–100High review priority

Multiple signals deserve documented manual review before reliance.

Published judgment

Every band change has a visible reason.

Check input coverage before interpreting any result.

Strongest aggravating signalBase band

High severity starts high, moderate starts moderate, and low starts low. Information-only and neutral observations do not raise priority.

Independent cause familiesAt most +1 band

Two or more unrelated causes raise the starting band once. Repeated views of the same cause are shown but do not count as corroboration.

Supporting evidenceAt most −1 band

Moderate-or-stronger supporting evidence can offset corroboration or lower a non-high starting band. It cannot erase an unrelated high-severity signal.

Within-band ordering0–100

After the band is fixed, the number orders documents within that band: 40% peak severity, 35% distinct cause families (capped at three), and 25% aggravating-signal volume (capped at five). It never chooses the band.

Signal catalog

Severity is published; points are not assigned per signal.

HighObserved evidence

Binary-format mismatch; metadata that names a document-generator tool.

ModerateObserved evidence

General-purpose editor marker; reversed timestamps; embedded script or attachment; a generic generator pipeline when a covered institution is also claimed.

LowObserved evidence

Ordinary timestamp gap; incremental revision chain; a generic generator pipeline without an institution claim.

Context onlyObserved evidence

Image evidence limits, EXIF presence or absence, software presence or absence, digital-signature containers, and an institution named in the file.

Interpretation standard

Observation, context, next evidence.

Each finding describes what the file exposes, names common legitimate causes, and recommends the next evidence that could resolve the question. Ocolta does not collapse missing evidence into a negative verdict.

Ruleset status

Version 2.0 is a transparent heuristic triage ruleset, not a validated fraud-detection benchmark. No accuracy, false-positive, or false-negative claim is made until a labeled dataset and published test methodology exist.

Deep Review methodology

A structured model review—not ruleset 2.0.

Deep Review is non-deterministic and may produce different or incomplete observations. It must keep fraud-risk indicators separate from AI-generation indicators and attach a page or region, observation, confidence, why the indicator may matter, a plausible benign explanation, and a proportionate next step.

Confidence applies to the reported observation. It is not a fraud probability, authenticity score, or statement about a person's intent. If the file is unreadable or materially incomplete, the appropriate result is unable to determine—not a forced conclusion.

No AI accuracy shortcut.

Deep Review is not presented as a fraud verdict or definitive AI-generation detector. Model output is supporting evidence for human review and source verification, never proof.

Open Deep Review or compare its processing boundary.

Institution fingerprint program

Twenty institutions, published traits, dated review.

Ocolta maintains two fingerprint databases. The genuine-format database records publicly observable conventions of authentic documents for the 20 institutions the guide library covers—12 banks (Chase, Bank of America, Wells Fargo, Citi, U.S. Bank, PNC, Truist, Capital One, TD Bank, Regions, Chime, Navy Federal) and 8 payroll providers (ADP, Gusto, Paychex, Workday, QuickBooks Payroll, Paycom, Paylocity, TriNet). Every trait is drawn from the published guide for that institution, so the database and the guides cannot drift apart, and every trait is hedged with “typically” or “commonly” because issuers change formats without notice.

The generator-signature database is deterministic and matches only PDF Producer and Creator metadata. It has exactly two confidence levels. High is reserved for metadata that names a document-generator tool in its own words—a Producer string containing “paystub maker” or “check stub generator”. Medium covers general-purpose HTML-to-PDF stacks (wkhtmltopdf, dompdf, TCPDF, mPDF, FPDF, jsPDF, WeasyPrint, PhantomJS, and the Chromium print pipeline). Template-generator sites commonly render with these stacks, and so do a great many legitimate businesses, so a medium match is always worded as a provenance question and never as an accusation.

What a fingerprint match is not.

PDF metadata can be edited, stripped, or inherited from honest reprocessing, and a document detected as naming an institution has only made a claim about itself. A fingerprint match is not proof of fraud, and its absence is not proof of authenticity. Reviewer uses the page they already have. Deep Review stays opt-in. The result stands without it.

Maintenance: fingerprint entries carry a version and a last-reviewed date, and are revised whenever the backing guide is revised. Where an institution has no single house format—Workday payslips are configured per employer, QuickBooks Payroll stubs are often deliberately plain—the entry says so, so that ordinary variation is not mistaken for a signal.

Deep Review+ corroboration

Research is constrained to institution-typed fields.

Deep Review+ adds one bounded research pass after the document review is complete. The model extracts institution-level entities from the document—employer, bank or statement issuer, payroll provider—and each one is checked in up to four ways: a public web-presence search for the organization name; a search associating the organization’s printed business address with its name; a public domain-registration (RDAP) lookup when the document prints the organization’s web domain; and a payroll ACH descriptor comparison against a published table of provider conventions.

Each entity is reported as corroborated, partially corroborated, not found, conflicting, or not researchable, with cited source links, a confidence level, and its own limitations. “Not found” means the automated pass located nothing—an inconclusive outcome. Small, new, family-run, and offline businesses routinely have no web footprint, and Ocolta will not convert that absence into an adverse inference.

Guardrails and limits.

Ocolta drops person-typed entities, rebuilds every query from allowlisted institution fields, and screens sensitive patterns before sending it. The permitted fields are an institution name and, when printed, its business address or domain. Automated extraction and classification can still be wrong, so users should remove personal data the review does not need. Ocolta is not a consumer reporting agency, does not investigate individuals, and does not produce consumer reports.

Bounds: at most four entities and six searches per review, a search deadline independent of the model deadline, and a hard rule that a search-provider failure degrades the appendix to an honest “not researchable” rather than voiding a paid review. The document itself is never sent to the search provider.

Start with evidence

Run the rules against a safe altered sample.

Run a private, browser-based integrity scan. Your file never leaves your device.

Scan a document