How Deft scores trust.

What each score means, how coordinated activity is flagged, and where the method stops. If a conclusion in a Deft report can't be traced back to this page, treat it as a bug.

Methodology v0.1Updated 5 Oct 2026

Relevance gate

Before anything is scored, each post has to be about the subject. Every study starts with a subject card: the subject's name, aliases, known namesakes and what is in scope. A relevance model reads each post against that card and gives it one of three labels.

LabelWhat happens
About the subjectGoes on to scoring and clustering.
UnrelatedSet aside. Typically a namesake, such as another business with the same name.
Not enough contextSet aside rather than guessed. A short reply with no names in it can't be attributed honestly.

The gate is the cheapest place to stop a wrong conclusion: a complaint about a namesake that slips through would otherwise be scored, clustered and reported as if it were about your subject. We check the gate against posts labelled by people, including deliberately confusing namesakes, and report results per subject rather than as a single average.

Three separate scores

Every record that passes the gate carries three scores. They are stored in separate columns and shown side by side. Nothing in the pipeline merges them into a single number, because a single number hides the reason behind a conclusion.

ScoreRangeWhat it answersHow it's computed
Source reliability0 to 1How much weight should this source carry?Merits minus demerits for the collection lane, the domain and provenance.
Output validity0 to 1Is the extracted record sound?A schema check plus a groundedness check that the extracted claim matches a span of the source text.
Sentiment−1 to +1What tone is expressed, and how strongly?Lexicon tone in English and Bahasa Melayu, plus an intensity score from 0 to 1. A language model can be used instead where configured.

Validity and sentiment are fixed once a record is scored. Reliability is the only score that can change later, and only for the reason described below.

Source reliability

Reliability describes the source, not the opinion. A verified outlet and a brand-new throwaway account can say the same thing; Deft weighs them differently. The score starts from where and how the record was collected, then adjusts for the domain and its provenance.

When a source dominates a cluster that gets flagged as coordinated, its reliability is reduced by 0.25 for later analysis. This is how one confirmed pattern makes the next one easier to see.

Coordinated activity

Records are embedded and grouped by text similarity. Each group, or cluster, is then checked against three conditions. A cluster is flagged only when all three hold:

  1. Size: at least 4 records.
  2. Intensity: average sentiment intensity of 0.6 or more.
  3. Concentration: at least 80% of the records come from one source, or from one day.

A flagged cluster gets an astroturf score, which reports how little the sources behind it can be trusted:

astroturf score = 1 − weighted reliability of the cluster's records

The sign of the cluster's average sentiment tells you its direction: manufactured praise when positive, a coordinated attack when negative. Thresholds are defaults and can be tuned per study; reports state the values that were used.

A high astroturf score means the activity looks coordinated and comes from weak sources. It does not tell you who organised it or why. Treat it as a reason to look closer, with the records in hand.

Early-warning triggers

Deft tracks clusters over time. A topic is promoted to an established category only after its volume has burst at least twice; one-day blips stay provisional.

It also looks for sources whose negativity has historically risen a number of days before another's. If that precursor is elevated now, Deft raises an early warning and reports its confidence. This is correlation, not proof, and confidence improves as a study builds history.

How a run is bounded

Collection is planned before it happens. A study's goal and policy produce a collection plan; Deft estimates its cost; a person approves that exact plan; only then does a run start. Every run is tied to the approved plan and logged.

Each workspace has a spend cap. A run that would exceed it stops before the paid call is made, not after.

Limits and human review

  • Public data shows what was said, not what is true. Reports separate what was observed from what was inferred, and state what remains uncertain.
  • A flag is evidence for a person to review. Deft does not accuse individuals and does not decide intent.
  • Small studies produce weaker signals. Where a cluster is close to a threshold, the report says so.
  • Sentiment lexicons cover English and Bahasa Melayu. Heavy slang, sarcasm and mixed-language posts are harder to score. We have not yet published an accuracy benchmark for Malay or mixed-language text.

Data handling

  • Public sources only. No private accounts and no logged-in collection. The crawler identifies itself as DeftCrawlerBot and honours robots.txt.
  • Pseudonymised authors. Author handles are salt-hashed before storage. Reports show stable pseudonyms, and re-identifying authors is not permitted.
  • Isolated workspaces. Each client workspace is isolated at the database level, so a member can read only their own workspace's data.
  • Retention. Study data is kept for the retention period set for the workspace.

Deft's formal privacy notice is in legal review. Questions about personal data can be sent to [email protected].