← Kabar Kopi

Content analysis and news clustering methodology

Kabar Kopi uses an AI-assisted content-analysis workflow to group coffee-industry news. It draws on research in content analysis, thematic analysis, semantic text representation, and classification evaluation. This is an operational editorial method; it is not a claim that Kabar Kopi conducts formal academic research or has independently validated every label.

How articles are read and coded

  1. Unit of analysis and extraction. A publisher article is the unit of analysis. The system records its URL, source, date, extraction status, headline, and available context. A headline or feed snippet is not treated as the full article. Thin or unavailable context is held for review or a retrieval retry.
  2. Relevance and headline–body alignment. The article context is used to decide whether coffee is central or incidental and whether the body supports the headline’s focus. Mismatch or uncertainty prevents automatic assignment.
  3. Paraphrased thematic code. The system writes one sentence in its own words to capture the central proposition: actor or subject, action or change, object, and supported impact. This is an analytical code, not a direct quotation. A separate verbatim source passage is retained as audit evidence so an editor can trace the interpretation back to the article.
  4. Cluster mapping. The thematic code and article context are compared with each cluster’s scope, exclusions, and previously labeled examples. A shared word alone is insufficient. One primary cluster supports navigation; secondary themes may be tags. Items awaiting review or lacking enough context are recorded separately as “Unclassified”. “Other” is assigned only after an editor confirms that an article is relevant but does not fit the available clusters.
  5. New topics and evaluation. Unmatched articles may form candidate topics for review; candidates do not automatically become public categories. Quality should be measured on an editor-labeled sample across clusters using precision, recall, F1, abstention rate, error audits, and inter-coder agreement.

What “research-informed” means

The workflow uses explainable and auditable steps: define the unit, document coding rules and cluster scopes, separate thematic interpretation from textual evidence, allow abstention, and evaluate results against editor labels. Scientific literature informs these choices; it does not prove that every automated label is correct. Performance must be measured on Kabar Kopi’s own data.

Research references

Limits and corrections

A paraphrase can misstate an article’s focus, extraction can omit context, and language models or embeddings can conflate neighboring themes. Clusters describe coverage patterns in the feed; they do not verify claims or measure economic impact or public opinion. Readers can inspect source links, and editors can correct assignments. Explore news and topic maps.