KMeans + LLM Labels + Laya¶
This walkthrough runs
examples/kmeans_genlbs_laya.py,
a lightweight pipeline that does not need BERTopic. It clusters the same 1,500 AG News articles as
the BERTopic example with plain KMeans, asks a local LLM (Qwen3-1.7B) to
write a checkable label for every cluster, and then lets Laya verify each article against those
labels.
Where the BERTopic example describes topics by five keywords, this one describes them with an
LLM-written title and a one-sentence description, built with the bundled
reduce and classes prompts.
Run it¶
From the repository root:
The first run downloads the dataset slice and three models (the embedder, Qwen/Qwen3-1.7B and
Laya). Results are written to output/<hash>/; later runs with the same settings reuse them.
Segmentation fault on macOS
On macOS, the script can crash with Segmentation fault: 11 (or
OMP: Error #15: Initializing libomp.dylib, but found libomp.dylib already initialized) at
embedder.index.search(...) in step 4. FAISS and PyTorch each bundle their own copy of the
OpenMP runtime, and FAISS's multi-threaded search crashes when both are loaded. Run with a
single OpenMP thread:
The LLM and Laya still run on the GPU (MPS); only CPU-side OpenMP work is single-threaded. The BERTopic example is not affected, because it never searches the index.
Step by step¶
1–2. Data and embeddings¶
The same as in the BERTopic example: the first 1,500 AG News training articles, embedded with
nomic-embed-text-v1.5 and cached in a FAISS index (embeds.bin) by
LocalEmbedder.
3. Cluster with KMeans¶
kmeans = KMeans(n_clusters=N_TOPICS, random_state=RND_SEED)
cluster_labels = kmeans.fit_predict(embeds)
docs_df["topic_id"] = cluster_labels
4. Label every cluster with an LLM¶
For each cluster, the script averages its embeddings into a centroid, L2-normalizes it, and asks the embedder's FAISS index for the five nearest articles. Because the vectors are normalized, the index's inner product is the cosine similarity:
centroids = np.array([embeds[docs_df["topic_id"] == k].mean(axis=0) for k in range(N_TOPICS)])
norms = np.linalg.norm(centroids, axis=1, keepdims=True)
norms[norms == 0] = 1.0
normalized_centroids = (centroids / norms).astype(np.float32)
_, top_5_indices = embedder.index.search(normalized_centroids, 5)
embedder.close() # free the embedder before loading the LLM
The five articles go to the reduce prompt,
which names the cluster's dominant theme with a title, a summary and five keywords. That result
goes to the classes prompt,
which turns it into a short, checkable class title and a one-sentence description. The titles of
earlier clusters are passed as used_titles, so the LLM can avoid reusing them (simplified from
the script, which also falls back to the reduce title if a call fails):
gen = LocalGenerator(LLM_MODEL, json_mode=True, temperature=0.1, max_tokens=512)
with gen:
for topic_id in range(N_TOPICS):
top_docs = [docs[idx] for idx in top_5_indices[topic_id]]
reduce_out = gen.batch(
reduce_tpl,
[{"text": json.dumps({"facts": top_docs}, indent=2)}],
req_keys=["title", "summary", "keywords"],
)[0]
class_out = gen.batch(
classes_tpl,
[{"text": json.dumps(reduce_out, indent=2), "used_titles": used_titles_str}],
req_keys=["title", "desc"],
)[0]
used_titles.append(class_out["title"])
LocalGenerator retries an answer until it parses as JSON with the required keys (see
JSON mode and retries), and turns
Qwen3's thinking mode off, so the 512-token budget goes to the answer.
5. Verify with Laya¶
Each topic becomes a Laya criterion that maps the class title to its description and keywords:
criteria_text = f"{row['Description']} (Keywords: {', '.join(kw_list)})"
topic_criteria.append(json.dumps({row["Name"]: criteria_text}))
results = jev.entail(docs=docs, lbls=topic_criteria)
6. Agreement per topic¶
As in the BERTopic example: the share of each cluster's articles that Laya assigns back to it.
Reading the results¶
One run on an Apple Silicon Mac (2 min 39 s with the models already downloaded) produced these labels and agreement rates; exact numbers and generated labels vary with hardware and library versions.
| Cluster | Generated class title | Articles | Agreement |
|---|---|---|---|
| 0 | Baseball Game Outcomes and Performances | 182 | 82.42% |
| 1 | Oil Price Drop-Induced Stock Market Recovery | 163 | 87.73% |
| 2 | Ocean Dead Zone Monitoring | 280 | 26.07% |
| 3 | Peace Bid in Najaf | 277 | 33.94% |
| 4 | Regulatory Approval for Public Offering | 195 | 26.67% |
| 5 | Windows Security Patching Tools | 290 | 46.55% |
| 6 | U.S. Swimming Performance in Olympic Competition | 113 | 22.12% |
Console output (descriptions shortened)
--- Discovered Topics ---
Topic 0 (n=182): Baseball Game Outcomes and Performances - Baseball game outcomes and performances are examined, ...
Topic 1 (n=163): Oil Price Drop-Induced Stock Market Recovery - Oil price drops drive U.S. stock market recovery ...
Topic 2 (n=280): Ocean Dead Zone Monitoring - Ocean dead zones are studied through advanced environmental monitoring ...
Topic 3 (n=277): Peace Bid in Najaf - The Iraqi Peace Mission in Najaf aims to end a radical Shi'ite uprising ...
Topic 4 (n=195): Regulatory Approval for Public Offering - Regulatory approval is sought in the final stages of Google's IPO ...
Topic 5 (n=290): Windows Security Patching Tools - Windows security patching tools are used to address vulnerabilities ...
Topic 6 (n=113): U.S. Swimming Performance in Olympic Competition - The U.S. swimming team faced challenges in key events ...
--- Topic Alignment & Agreement ---
Topic 0 [Baseball Game Outcomes and Performances]: Agreement between KMeans and Laya: 82.42%
Topic 1 [Oil Price Drop-Induced Stock Market Recovery]: Agreement between KMeans and Laya: 87.73%
Topic 2 [Ocean Dead Zone Monitoring]: Agreement between KMeans and Laya: 26.07%
Topic 3 [Peace Bid in Najaf]: Agreement between KMeans and Laya: 33.94%
Topic 4 [Regulatory Approval for Public Offering]: Agreement between KMeans and Laya: 26.67%
Topic 5 [Windows Security Patching Tools]: Agreement between KMeans and Laya: 46.55%
Topic 6 [U.S. Swimming Performance in Olympic Competition]: Agreement between KMeans and Laya: 22.12%
Where the articles went (rows: KMeans clusters; columns: Laya's classes):
| Cluster ↓ · class → | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 Other | All |
|---|---|---|---|---|---|---|---|---|---|
| 0 Baseball | 150 | 0 | 0 | 1 | 0 | 0 | 8 | 23 | 182 |
| 1 Oil prices, stocks | 1 | 143 | 1 | 2 | 0 | 0 | 0 | 16 | 163 |
| 2 Ocean dead zones | 10 | 3 | 73 | 5 | 0 | 5 | 0 | 184 | 280 |
| 3 Najaf | 18 | 9 | 2 | 94 | 0 | 3 | 0 | 151 | 277 |
| 4 Public offering | 13 | 18 | 0 | 1 | 52 | 7 | 0 | 104 | 195 |
| 5 Windows security | 7 | 2 | 0 | 2 | 10 | 135 | 0 | 134 | 290 |
| 6 Olympic swimming | 57 | 0 | 0 | 5 | 0 | 1 | 25 | 25 | 113 |
| All | 256 | 175 | 76 | 110 | 62 | 151 | 33 | 637 | 1500 |
What this run suggests:
- Precise labels, narrow coverage. Each label is written from only the five articles nearest the centroid, so it describes the center of the cluster, not all of it. Cluster 2 is mostly Sci/Tech (204 of 280 by the gold labels), but only 4% of its articles mention oceans or dead zones; Laya keeps 73 and sends 184 to Other. Likewise, a quarter of cluster 4 mentions Google or an IPO, against 81% of the 52 articles Laya keeps.
- Clear themes hold up. Baseball (82%) and oil prices and stocks (88%) are confirmed for most of their articles.
- Neighbouring labels compete. Cluster 6 is about the Olympics (81% of its articles mention them), but its label is about swimming, which only 28% mention. Laya moves 57 of its articles to the baseball label, whose description ("playoff results, tour victories") reads as general sports.
- More Other than with keywords. 637 articles (42%) match no label, against 449 (30%) in the BERTopic example. The two runs cluster differently (KMeans on the raw embeddings here, UMAP + KMeans there), so this is not a like-for-like comparison of the labeling methods.
The trade-off is the one described in How It Works: a
precise, checkable label confirms a tight core and flags the rest. To cover more of each cluster,
feed reduce more than the five documents nearest the centroid, for example a larger sample of
compressed documents.
Note
Small models do not always follow the finer prompt rules: "Baseball Game Outcomes and
Performances" joins two ideas with "and", which the classes prompt forbids. See
Prompt Templates.
Full script¶
examples/kmeans_genlbs_laya.py
import json
import hashlib
from pathlib import Path
import numpy as np
import pandas as pd
from dotenv import load_dotenv
from datasets import load_dataset
from sklearn.cluster import KMeans
from topicjev.entail import LayaEntail
from topicjev.gen import LocalEmbedder, LocalGenerator
from topicjev.prompts import load_prompt
load_dotenv()
# Dataset and model configurations
DATASET_NAME: str = "fancyzhx/ag_news"
N_SAMPLES: int = 1500
N_TOPICS: int = 7
RND_SEED: int = 42
EMBED_MODEL: str = "nomic-ai/nomic-embed-text-v1.5"
EMBED_PREFIX: str = "clustering: "
LLM_MODEL: str = "Qwen/Qwen3-1.7B"
JEV_MODEL: str = "convaiinnovations/laya-typed-decisions"
# Output directory keyed by CONFIG for reproducible caching
CONFIG = {
"dataset": DATASET_NAME,
"n_samples": N_SAMPLES,
"embed_model": EMBED_MODEL,
"embed_prefix": EMBED_PREFIX,
"llm_model": LLM_MODEL,
"jev_model": JEV_MODEL,
"n_topics": N_TOPICS,
"seed": RND_SEED,
}
OUT_DIR: Path = Path("./output") / hashlib.sha1(json.dumps(CONFIG).encode()).hexdigest()[:8]
OUT_DIR.mkdir(parents=True, exist_ok=True)
(OUT_DIR / "config.json").write_text(json.dumps(CONFIG, indent=2))
print(f"Writing results to {OUT_DIR}")
# 1. Load dataset (1500 samples from ag_news)
if not (OUT_DIR / "docs.csv").exists():
ds = load_dataset(DATASET_NAME, split=f"train[:{N_SAMPLES}]")
docs_df = ds.to_pandas()
docs_df["id"] = range(len(docs_df))
docs_df.to_csv(OUT_DIR / "docs.csv", index=False)
docs_df = pd.read_csv(OUT_DIR / "docs.csv")
docs = docs_df["text"].tolist()
# 2. Compute dense embeddings with FAISS index
embedder = LocalEmbedder(
EMBED_MODEL,
prefix=EMBED_PREFIX,
trust_remote_code=False,
save_dir=str(OUT_DIR),
)
embeds = embedder.get_embeds(docs)
# 3. Cluster using only KMeans
if not (OUT_DIR / "doc_labels.csv").exists():
kmeans = KMeans(n_clusters=N_TOPICS, random_state=RND_SEED)
cluster_labels = kmeans.fit_predict(embeds)
docs_df["topic_id"] = cluster_labels
docs_df.to_csv(OUT_DIR / "doc_labels.csv", index=False)
docs_df = pd.read_csv(OUT_DIR / "doc_labels.csv")
# 4. Generate Topic Labels (Reduce Prompt -> Classes Prompt) with Qwen3-1.7B
if not (OUT_DIR / "topic_info.csv").exists():
# Derive cluster centroids directly from embeddings and assigned topic labels
centroids = np.array([
embeds[docs_df["topic_id"] == k].mean(axis=0)
for k in range(N_TOPICS)
])
# L2-normalize centroids so inner product equals cosine similarity
norms = np.linalg.norm(centroids, axis=1, keepdims=True)
norms[norms == 0] = 1.0
normalized_centroids = (centroids / norms).astype(np.float32)
_, top_5_indices = embedder.index.search(normalized_centroids, 5)
# Free embedding model from memory before loading LLM and entailment models
embedder.close()
reduce_tpl = load_prompt("reduce")
classes_tpl = load_prompt("classes")
gen = LocalGenerator(
LLM_MODEL,
json_mode=True,
temperature=0.1,
max_tokens=512,
)
topic_records = []
used_titles: list[str] = []
with gen:
for topic_id in range(N_TOPICS):
# Pull top 5 representative documents from FAISS
top_docs = [docs[idx] for idx in top_5_indices[topic_id]]
facts_json = json.dumps({"facts": top_docs}, indent=2)
# Step 4a: Run reduce prompt to synthesize dominant theme
print(f"\n[Topic {topic_id}] Generating synthesis via reduce prompt...")
reduce_results = gen.batch(
reduce_tpl,
[{"text": facts_json}],
req_keys=["title", "summary", "keywords"],
)
reduce_out = reduce_results[0]
print(f" Reduce Title: {reduce_out.get('title')}")
print(f" Keywords: {reduce_out.get('keywords')}")
# Step 4b: Pass reduce output into classes prompt to derive an entailment-checkable class
print(f"[Topic {topic_id}] Generating checkable class label via classes prompt...")
input_json = json.dumps(reduce_out, indent=2)
used_titles_str = json.dumps(used_titles) if used_titles else "None"
class_results = gen.batch(
classes_tpl,
[{"text": input_json, "used_titles": used_titles_str}],
req_keys=["title", "desc"],
)
class_out = class_results[0]
final_title = class_out.get("title", reduce_out.get("title", f"Topic_{topic_id}"))
final_desc = class_out.get("desc", reduce_out.get("summary", ""))
used_titles.append(final_title)
print(f" Class Title: {final_title}")
print(f" Description: {final_desc}")
topic_records.append({
"Topic": topic_id,
"Count": int((docs_df["topic_id"] == topic_id).sum()),
"Name": final_title,
"Description": final_desc,
"Keywords": json.dumps(reduce_out.get("keywords", [])),
"Reduce_Title": reduce_out.get("title", ""),
"Summary": reduce_out.get("summary", ""),
"Representative_Docs": json.dumps(top_docs),
})
tinfo = pd.DataFrame(topic_records)
tinfo.to_csv(OUT_DIR / "topic_info.csv", index=False)
else:
# Free embedding model from memory if topic_info is already cached
embedder.close()
tinfo = pd.read_csv(OUT_DIR / "topic_info.csv")
print("\n--- Discovered Topics ---")
for _, row in tinfo.iterrows():
print(f"Topic {row['Topic']} (n={row['Count']}): {row['Name']} - {row['Description']}")
# 5. Bootstrap Entailment with Laya (JevLike model)
if not (OUT_DIR / "jev_labels.csv").exists():
jev = LayaEntail(JEV_MODEL)
with jev:
# Formulate criteria schema mapping each topic name to its description & keywords
topic_criteria = []
for _, row in tinfo.iterrows():
kw_list = json.loads(row["Keywords"]) if isinstance(row["Keywords"], str) else []
criteria_text = f"{row['Description']} (Keywords: {', '.join(kw_list)})" if kw_list else row["Description"]
topic_criteria.append(json.dumps({row["Name"]: criteria_text}))
results = jev.entail(docs=docs, lbls=topic_criteria)
docs_df["jev"] = [r["class"] for r in results]
docs_df.to_csv(OUT_DIR / "jev_labels.csv", index=False)
docs_df = pd.read_csv(OUT_DIR / "jev_labels.csv")
# 6. Calculate agreement between KMeans clusters and Laya entailment decisions
print("\n--- Topic Alignment & Agreement ---")
for topic_id in sorted(docs_df["topic_id"].unique()):
topic_docs = docs_df[docs_df["topic_id"] == topic_id]
agreement = (topic_docs["topic_id"] == topic_docs["jev"]).mean()
topic_name = tinfo.loc[tinfo["Topic"] == topic_id, "Name"].values[0]
print(f"Topic {topic_id} [{topic_name}]: Agreement between KMeans and Laya: {agreement:.2%}")