Zero-shot Classification¶
topicjev.entail decides which of a set of candidate labels a document belongs to, without any
training data. It is the core of the TopicJev bootstrap: the labels are the topics of a topic model,
and the classifier checks whether each document really fits its topic.
All backends share one method:
and differ only in how they score a (document, label) pair.
Results¶
entail() returns one dictionary per document, in input order:
| Key | Type | Meaning |
|---|---|---|
class |
int |
Index of the assigned label in lbls. len(lbls) means Other: no label passed the threshold, or a decoy won. |
prob |
float |
Probability of the top-scoring candidate, whether a label or a decoy. |
other |
str or None |
The decoy's text when a decoy scored highest, otherwise None. |
probs |
list[float] |
One probability per label in lbls, in the same order. Decoys are left out, so with decoys the values can sum to less than 1. |
To turn results into label names:
Choosing a backend¶
| Backend | Model family | Score of a (document, label) pair | Default temperature | Multi-label |
|---|---|---|---|---|
XEncoderEntail |
NLI cross-encoder (e.g. DeBERTa-v3) | Entailment logit of "The text discusses label" | 1.0 |
Yes |
GoalExEntail |
Encoder-decoder (e.g. FLAN-T5) | \(\log P(\text{yes}) - \log P(\text{no})\) at the first decoder step | 0.1 |
Yes |
CausalEntail |
Causal chat LM (e.g. Qwen, Llama) | \(\log P(\text{yes}) - \log P(\text{no})\) at the next token | 0.1 |
Yes |
Seq2SeqEntail |
Encoder-decoder (e.g. FLAN-T5) | Mean log-likelihood of the label's tokens as the answer | 1.0 |
Yes |
LayaEntail |
Local Laya decision model (ModernBERT) | Laya's choice probabilities | 1.0 |
No |
JevEntail |
Hosted TypeSafe API | TypeSafe System-1 choice probabilities | 1.0 |
No |
As a rule of thumb:
- Start with
XEncoderEntailorLayaEntail: they are small, fast and made for this task. - Use
GoalExEntailorCausalEntailwhen a label is a longer property or claim that benefits from an instruction-following model. - Local backends score every document against every label, so the cost grows with
len(docs) * len(labels). Laya and Jev score all labels of a document in one call.
Backends¶
XEncoderEntail¶
Uses a Natural Language Inference cross-encoder. Each document is the premise; each label is
turned into a hypothesis with a template, by default "The text discusses {}".
from topicjev.entail import XEncoderEntail
docs = [
"NASA's Artemis program plans lunar base camp and rover deployments.",
"Central banks consider interest rate cuts amid slowing inflation.",
]
labels = ["Space Exploration", "Monetary Policy", "Sports Organization"]
model = XEncoderEntail("MoritzLaurer/deberta-v3-base-zeroshot-v2.0")
with model:
results = model.entail(docs=docs, lbls=labels)
# results, rounded to 4 digits
[{'class': 0, 'prob': 0.9997, 'other': None, 'probs': [0.9997, 0.0001, 0.0002]},
{'class': 1, 'prob': 0.9998, 'other': None, 'probs': [0.0001, 0.9998, 0.0001]}]
| Parameter | Default | Meaning |
|---|---|---|
hyp |
"The text discusses {}" |
Hypothesis template; {} is replaced by the label. |
The positions of the entailment, contradiction (or not_entailment) and neutral outputs are
read from the model's id2label config, so both 2-class zero-shot models and 3-class NLI models
work. In single-label mode the raw entailment logits of all labels go through a softmax. In
multi-label mode, each label gets entailment minus contradiction (and
neutral) and then a sigmoid.
GoalExEntail¶
Asks an encoder-decoder model a Yes/No question per label and compares the scores of the Yes and No tokens at the first decoder step, in one forward pass and without generating text. The prompt follows the GoalEx property-verification format:
Check whether the TEXT satisfies a PROPERTY. Respond with Yes or No. When uncertain, output No.
Now complete the following example -
input:
PROPERTY: {property}
TEXT: {text}
output:
from topicjev.entail import GoalExEntail
model = GoalExEntail("google/flan-t5-base")
with model:
results = model.entail(docs=docs, lbls=labels)
| Parameter | Default | Meaning |
|---|---|---|
template |
the GoalEx prompt above | Prompt with {property} and {text} placeholders. |
reserve |
8 |
Tokens kept free when truncating the document to fit the context. |
temperature |
0.1 |
Divides the Yes/No scores before normalization. |
Every spelling of yes and no that the tokenizer encodes as a single token (yes, Yes,
YES, ...) is combined with logsumexp. A tokenizer without any single-token variant raises a
ValueError.
CausalEntail¶
The same Yes/No question as GoalExEntail, asked to a causal chat model. The prompt is wrapped in
the model's chat template, and the score is read from the logits of the next token.
from topicjev.entail import CausalEntail
model = CausalEntail("Qwen/Qwen2.5-1.5B-Instruct")
with model:
results = model.entail(docs=docs, lbls=labels)
| Parameter | Default | Meaning |
|---|---|---|
template |
the GoalEx prompt | Prompt with {property} and {text} placeholders. |
chat |
True |
Wrap the prompt in the tokenizer's chat template, if it has one. |
thinking |
False |
Keep "thinking" enabled for reasoning models (e.g. Qwen3); off by default, so the model answers right away. |
reserve |
8 |
Tokens kept free when truncating the document. |
temperature |
0.1 |
As for GoalExEntail. |
The tokenizer is left-padded, so the last position of every row in a batch is the end of the prompt.
Seq2SeqEntail¶
Scores each label by how likely the model is to produce it as the answer to
"Text: {document}\nThe topic of this text is:": the mean log-probability of the label's tokens
under teacher forcing.
from topicjev.entail import Seq2SeqEntail
model = Seq2SeqEntail("google/flan-t5-base")
with model:
results = model.entail(docs=docs, lbls=labels)
| Parameter | Default | Meaning |
|---|---|---|
template |
"Text: {}\nThe topic of this text is:" |
Prompt with one {} placeholder for the document. |
reserve |
50 |
Tokens kept free when truncating the document. |
LayaEntail and JevEntail¶
Laya is an open-source
ModernBERT decision model in the style of TypeSafe's "System One" models: it answers a
multiple-choice question about a text. LayaEntail runs it locally; JevEntail sends the same
question to the hosted TypeSafe Jev API.
from topicjev.entail import LayaEntail
jev = LayaEntail("convaiinnovations/laya-typed-decisions")
with jev:
results = jev.entail(docs=docs, lbls=labels)
# results, rounded to 4 digits
[{'class': 0, 'prob': 0.8579, 'other': None, 'probs': [0.8579, 0.0350, 0.0340]},
{'class': 1, 'prob': 0.8544, 'other': None, 'probs': [0.0382, 0.8544, 0.0385]}]
Labels with descriptions. A label can be a plain name, or a JSON object that maps the name to a description. Laya uses the description to decide; this is how the BERTopic example passes each topic's keywords:
import json
labels = [
json.dumps({"Space Exploration": "rockets, missions, astronauts, the Moon"}),
json.dumps({"Monetary Policy": "central banks, interest rates, inflation"}),
json.dumps({"Sports Organization": "leagues, clubs, tournaments"}),
]
Use one key per object, and keep label names unique.
A built-in Other option. Both classes add the choice
{"Other": "None of the other categories fit this text"} as a decoy. That is why the
probs above sum to about 0.93: the rest went to Other. When Other wins, class is
len(labels) and other holds that JSON string.
Always single-label. The model's choice probabilities are used as they are (with the default
temperature of 1.0), and multi_lbl is ignored.
Uncalibrated confidence
Confidence calibration has not been tuned for Laya or Jev. The decision is the top category
plus decoy routing; treat prob as an uncalibrated indicator, not as a posterior probability.
Laya itself warns about this when it loads the laya-typed-decisions checkpoint.
LayaEntail |
JevEntail |
|
|---|---|---|
| Runs | Locally, on the detected device | On the TypeSafe API |
| Constructor | LayaEntail(model_name, **options) |
JevEntail(**options) (no model name) |
| Batching | batch_size documents per call |
One request per document (large batches not yet evaluated) |
| Credentials | None | TYPESAFE_API_KEY in the environment or a .env file |
Scoring and calibration¶
Every backend goes through the same steps. For documents \(d\), labels \(l\) and decoys:
- Pair every document with every label and decoy, truncating the document to fit the model.
- Score each pair with the backend's raw score \(s(d, l)\).
- Calibrate with probes, if given: \(s(d, l) \leftarrow s(d, l) - \frac{1}{|P|}\sum_{p \in P} s(p, l)\).
- Scale by the temperature \(T\): \(s(d, l) / T\).
- Normalize: a softmax across labels and decoys, or, in multi-label mode, a sigmoid per label.
- Decide: take the top candidate. It becomes the
classonly if it is a real label and its probability is above the threshold; otherwise the document is Other.
All options below are constructor arguments and work with every backend.
Decoys¶
Under a closed list of labels, a softmax always picks one of them, even for a document that fits
none. Decoys are extra, plausible labels that compete for probability. If a decoy wins, the
document is assigned to Other, and other says which decoy took it. With the three labels above,
the text "Heavy rain and thunderstorms expected over the weekend." comes back as:
Probes¶
model = CausalEntail("Qwen/Qwen2.5-1.5B-Instruct", probes=["", "N/A", "Lorem ipsum dolor sit amet."])
Some labels score high for almost any text: the model is biased towards them. Probes are content-free texts; their mean score per label is subtracted from every document's score before normalization, which removes that bias (contextual calibration).
Temperature¶
scores / temperature before the softmax or sigmoid. A lower temperature sharpens the
distribution, a higher one flattens it. GoalExEntail and CausalEntail default to 0.1, the
other backends to 1.0.
Tip
With the default 0.1, the Yes/No backends often return probabilities of almost exactly 0 or 1
(in our tests, CausalEntail with Qwen2.5-1.5B gave prob=1.0 on clear-cut examples). That
is fine for picking a label. If you want to rank documents by confidence or compare against a
threshold, raise the temperature, e.g. temperature=1.0.
Threshold¶
The top label is accepted only if its probability is above the threshold. The default depends on the mode:
- single-label: \(2 / N\) for \(N\) labels, i.e. twice the uniform probability;
- multi-label: \(0.5\).
With one or two labels, the default threshold rejects everything
For \(N = 2\) the default is \(2/2 = 1.0\), and no probability is ever above it, so every
document becomes Other, even at 99.99% confidence. Pass an explicit value, e.g.
threshold=0.5. To check a single label, use multi_lbl=True instead: a softmax over one
label is always 1.0, while the sigmoid of multi-label mode judges the label on its own.
Multi-label mode¶
With multi_lbl=True, every label is scored independently with a sigmoid, so several labels can
have high probability at once. class still reports only the best label; read probs to get all
labels that apply:
LayaEntail and JevEntail are always single-label.
Common parameters¶
| Parameter | Default | Meaning |
|---|---|---|
batch_size |
16 |
Pairs per forward pass (documents per call for Laya). |
max_tokens |
512 |
Context length to use when neither the tokenizer nor the model config states one. |
multi_lbl |
False |
Sigmoid per label instead of a softmax across labels. |
decoys |
None |
Extra labels that send a document to Other when they win. |
probes |
None |
Content-free texts used to subtract label bias. |
temperature |
1.0 (0.1 for GoalEx and Causal) |
Divides the scores before normalization. |
threshold |
2 / N (single-label) or 0.5 (multi-label) |
Minimum probability to accept the top label. |
Long documents¶
Documents are truncated so that the prompt, the longest label and the answer fit the model's
context length. That length comes from the tokenizer (model_max_length) or the model config, and
falls back to max_tokens. Most cross-encoders accept 512 tokens, so a long document is judged by
its beginning only. Compress long documents first if their main point may come
later.
Writing your own backend¶
Subclass Entailment (or LocalEntailment for a PyTorch model, which adds device detection and
memory clean-up) and implement three methods:
load_model(): load weights or open a client;prep_pairs(prompts, lbls): build onePairper (document, label), documents in the outer loop;batch_score(pairs): return one raw score per pair (higher means more likely).
Probes, decoys, temperature, normalization and thresholds then work unchanged. A toy backend that counts shared words:
from topicjev.entail import Entailment, Pair
class KeywordEntail(Entailment):
"""Scores a label by how many of its words appear in the document."""
def load_model(self) -> None:
pass # nothing to load
def prep_pairs(self, prompts: list[str], lbls: list[str]) -> list[Pair]:
return [Pair(inpt=doc.lower(), targ=lbl.lower()) for doc in prompts for lbl in lbls]
def batch_score(self, pairs: list[Pair]) -> list[float]:
return [float(sum(w in p.inpt for w in p.targ.split())) for p in pairs]
model = KeywordEntail(None, temperature=0.5)
with model:
results = model.entail(
docs=["Interest rates and inflation worry central banks."],
lbls=["space rockets", "interest rates inflation", "football league"],
)
print(results)
# [{'class': 1, 'prob': 0.995..., 'other': None, 'probs': [0.0025, 0.995..., 0.0025]}]