Nothing reaches you unlabelled. Centragon is why.
Between our collectors and the nine modules sits one layer that every record passes through. It decides what a post is, who it is aimed at, how serious it is and which ATT&CK techniques it describes, all before you ever run a query. This is not a tenth module. It is why the other nine return fields instead of paragraphs.
One stop on the way to your query
Collection decides what we see. Enrichment decides what it means. The modules are the windows onto the result, and by the time you open one the reasoning has already happened.
RECORD PATH
stage 03 of 04This is the reason a search here returns a category and a technique rather than a wall of forum text: the expensive thinking is done once, on ingest, for every record, not per query, per analyst, and not by you.
A schema, published
These are the label sets a record is scored against. We publish the schema because a schema is checkable: you can hold us to a category list. What we never publish is a volume: how many records carry a label describes our collection, and the shape of the collection is the one thing that has to stay ours.
32
Content categories
access sale, victim announcement, combolist, stealer logs, exploit sale, carding, sim swap, recruitment…
27
ATT&CK techniques
mapped across 8 tactics, so a listing arrives already placed on the kill chain
16
Target sectors
who the post is aimed at, not who wrote it
11
Data types
credentials, payment card, government ID, source code, database…
10
PII entity types
detected so they can be counted and masked, never returned as values
7
IOC classes
emails · IPs · domains · CVEs · file hashes · crypto wallets · credit cards
7
Motivations
financial, ideological, and the five that are neither
5
Severity levels
plus a separate relevance score, because "severe" and "yours" are different questions
4
Sophistication levels
a scripted resale and a targeted operation should not read alike
A listing in, a record out
The same pass that runs on every collected post, on a synthetic and clearly labelled listing. The numbered chips are the point: each says which phrase produced which field, which is the difference between scraping a forum and reading one.
UNDERGROUND LISTING
synthetic · values maskedtarget: EU fintech · rev ~$40M1
extras: OTP seed included2
sample: 3
SOURCE & AUTHOR FIELDS EXIST ON THE REAL RECORD · NEVER SHOWN HERE
INTEL RECORD
1–4 WHICH PHRASE PRODUCED WHICH FIELD · SAMPLE IS SYNTHETIC · DETECTED VALUES NEVER LEAVE THE PLATFORM
Five things this layer computes and keeps
A layer that reads the underground knows more than it should hand back. This is the list of what it deliberately does not, and it is the honest reason this page is not selling you access to a module.
Which source the post came from
The layer records it, and it is the single most sensitive field we hold: a source name is a collection posture, and publishing one closes the channel for everybody. Customers get the record and its date; nobody gets the address it came from.
The post's own words
Raw content and the model's prose summary stay inside. What crosses the boundary is the label set (category, sector, severity, techniques) because a structured field can be filtered, counted and trusted, while free text from a criminal forum is an injection surface wearing a paragraph.
The values behind the entity counts
A record will tell you it carried three email addresses, one IP and a card number. It will not tell you which. Detection exists so the material can be counted and masked; returning the values would make us the distributor of exactly what we monitor.
Who the post names
Target organizations, target countries and actor aliases are extracted for scoring and then withheld from public surfaces. An entitled organization sees its own exposure; nobody browses a list of other people's victims.
Our own footprint
Processing logs and blacklists say what we watch, how often, and what we gave up on. That is a map of the collection, so it is not a public field, and it is not one we trade for a demo either.
The labels, on their own meter
The layer is reachable directly, so your pipeline can ask for the labels instead of re-deriving them. Its ceilings are deliberately the lowest on the platform: a call here runs a model, not an index lookup, and pricing that pretends otherwise would be a bill you did not agree to.
Ceilings are daily, organization-wide and API-only; console work is unmetered. Past the ceiling further calls are declined, never an overage line. The figure per tier sits on the pricing page, beside the tier it belongs to, and is written down nowhere else.
Ask what a week of your exposure looks like, labelled.
Bring a domain. We run the pass live and show you the fields.
NDA-friendly briefings · global coverage · no slideware