Prompt Templates¶
topicjev.prompts ships prompt templates for describing topics with an LLM. They are plain text
files inside the package, loaded by name:
from topicjev.prompts import list_prompts, load_prompt
list_prompts() # ['classes', 'reduce']
template = load_prompt("reduce")
Both templates are format strings for LocalGenerator.batch(), and both ask for a
bare JSON object.
reduce: from facts to a theme¶
Input: short factual snippets from the documents of one cluster, for example the compressed documents, or the documents nearest the cluster's centroid. Output: the one dominant theme of the snippets, as a title, a literature-review style summary, and exactly five keywords. The prompt tells the model to ignore snippets that do not fit the main theme instead of stitching every snippet into the summary.
| Placeholder | Fill with |
|---|---|
{text} |
A JSON object {"facts": ["...", "..."]} |
Output keys: title, summary, keywords.
classes: from a theme to a checkable class name¶
Input: the output of reduce, plus the titles already used by other topics. Output: a 2–5 word
class name and a one-sentence description. The title is designed to complete the sentence
"The text discusses ______", which is exactly the default hypothesis of
XEncoderEntail. The prompt therefore asks for a single, concrete,
checkable claim: never two topics joined by "and", never a generic umbrella term, and distinct from
every used title.
| Placeholder | Fill with |
|---|---|
{text} |
The reduce result, as JSON |
{used_titles} |
Titles already given to other topics, one per line (can be empty) |
Output keys: title, desc.
Example: name a cluster, then check it¶
The KMeans + LLM labels example runs this chain on real
clusters: it feeds the five documents nearest each centroid to reduce, passes the result to
classes, and hands the titles and descriptions to LayaEntail. The steps in isolation:
import json
from topicjev.gen import LocalGenerator
from topicjev.prompts import load_prompt
facts = {
"facts": [
"Solid-state lithium cells with ceramic electrolytes resist dendrite growth.",
"A sulfide electrolyte raised ionic conductivity at room temperature.",
"Ceramic separators kept 92% capacity after 1,000 cycles.",
"A survey reports rising cobalt prices.",
]
}
gen = LocalGenerator("Qwen/Qwen2.5-1.5B-Instruct", temperature=0.0)
with gen:
theme = gen.batch(
load_prompt("reduce"),
[{"text": json.dumps(facts)}],
req_keys=["title", "summary", "keywords"],
)[0]
label = gen.batch(
load_prompt("classes"),
[{"text": json.dumps(theme), "used_titles": "Bus Fleet Electrification"}],
req_keys=["title", "desc"],
)[0]
print(theme["title"]) # Electrolyte and Separator Improvements
print(label["title"]) # Lithium Cell Dendrite Growth Prevention
The class title can now be checked against every document of the cluster. With a single label,
use multi_lbl=True: the label is then judged on its own, with a sigmoid and a threshold of 0.5.
(In single-label mode, a softmax over one label always gives 1.0.)
from topicjev.entail import XEncoderEntail
cluster_docs = [
"Solid-state lithium cells with ceramic electrolytes resist dendrite growth.",
"Ceramic separators kept 92% capacity after 1,000 cycles.",
"A survey reports rising cobalt prices.",
]
with XEncoderEntail("MoritzLaurer/deberta-v3-base-zeroshot-v2.0", multi_lbl=True) as model:
results = model.entail(docs=cluster_docs, lbls=[label["title"]])
for doc, res in zip(cluster_docs, results):
print(res["class"], round(res["probs"][0], 4), doc)
0 0.9991 Solid-state lithium cells with ceramic electrolytes resist dendrite growth.
1 0.0002 Ceramic separators kept 92% capacity after 1,000 cycles.
1 0.0003 A survey reports rising cobalt prices.
Class 0 confirms the document; class 1 (= len(lbls)) is Other. Only the first document is
confirmed: the title is precise enough to reject the cobalt-price outlier, and also narrow enough
to miss the capacity result. That trade-off is what the bootstrap measures.
Check what small models return
In our tests a 1.5B model always returned valid JSON, but did not always follow the finer rules: the reduce
title above, "Electrolyte and Separator Improvements", joins two topics with "and", which the
prompt forbids. Review a sample of the outputs, or try a larger model.
Your own templates¶
Any string works as a template for LocalGenerator.batch(). Keep two rules in mind:
- Placeholders are filled with
str.format(**input), so every{name}must be a key of each input dictionary. - Literal braces must be doubled: write
{{"title": "..."}}to show the model a JSON example.