Skip to content

Compression

topicjev.compress shortens documents before they are classified, embedded or labeled. Small classifiers read only the first few hundred tokens of a text (see Long documents), and LLM prompts that summarize a whole cluster get expensive quickly. A compressed document keeps its main point within that budget.

All compressors take a list of strings and return a list of strings of the same length:

with compressor:
    short_docs = compressor.compress(docs)

Choosing a compressor

Compressor Mechanism Output length Good for
CausalCompressor Summary by a causal chat LM Up to ratio × prompt length, in tokens Modern instruction-tuned LLMs (Qwen, Llama, Mistral); the budget grows with the document.
CNNCompressor Summary by a summarization-tuned seq2seq model Up to ratio × the model's context length Checkpoints fine-tuned on news summarization, such as BART-CNN or PEGASUS.
Seq2SeqCompressor Instruction-prompted seq2seq summary Up to ratio × the model's context length General instruction-tuned encoder-decoders such as FLAN-T5.
LinguaCompressor Token pruning with LLMLingua-2, no generation About ratio of the original tokens Fast, extractive compression that keeps the original wording.

Parameters shared by all compressors:

Parameter Default Meaning
ratio 0.3 Output budget; its meaning depends on the compressor, see the table above.
batch_size 8 Documents per generation batch (not used by LinguaCompressor).
source_len per compressor Input length to use when the tokenizer and model config do not state one.

The examples below use these two documents:

long_docs = [
    "Researchers developed a solid-state lithium battery with a ceramic electrolyte that resists dendrite growth. "
    "Over 1,000 charge cycles the cell retained 92 percent of its capacity, and it operated safely at temperatures "
    "up to 80 degrees Celsius. The authors argue the design could shorten charging times for electric vehicles, "
    "although manufacturing the thin ceramic layers at scale remains expensive.",
    "A city council approved a plan to electrify its entire bus fleet by 2030. The plan replaces 400 diesel buses, "
    "adds charging depots at three garages, and is funded partly by a federal grant. Officials expect lower "
    "maintenance costs and cleaner air along busy corridors, while unions asked for retraining programs for mechanics.",
]

CausalCompressor

Prompts a causal chat model to summarize each document. The token budget is relative to the input: max_new_tokens = ratio × prompt length, where the prompt length is the longest prompt in the batch, including the instruction and chat template.

from topicjev.compress import CausalCompressor

compressor = CausalCompressor("Qwen/Qwen2.5-1.5B-Instruct", ratio=0.3, batch_size=4)
with compressor:
    summaries = compressor.compress(long_docs)
Researchers created a solid-state lithium battery with a ceramic electrolyte that prevents
dendrite growth and maintains over 92% capacity after 1,000 charge cycles. It operates safely up to

The budget is a hard cut

Generation stops when the budget runs out, even mid-sentence, as in the output above. Raise ratio, or set min_len, if summaries come back truncated.

Parameter Default Meaning
instruct see below Prompt template with a {text} placeholder.
chat True Wrap the prompt in the tokenizer's chat template.
thinking False Keep "thinking" enabled for reasoning models.
min_len 0 Minimum number of new tokens; also a floor for the budget.
num_beams 1 Beam search width (1 = greedy decoding).
source_len 2048 Fallback input length.

The default instruction:

Summarize the text below in several declarative sentences in paragraph form. Output only the summary, with no preamble.

TEXT:
{text}

SUMMARY:

CNNCompressor

For seq2seq checkpoints fine-tuned to summarize, such as facebook/bart-large-cnn. The document goes in as is, without an instruction, and the summary is generated with beam search. The budget is fixed per model: max_length = ratio × context length (1024 tokens for BART).

from topicjev.compress import CNNCompressor

compressor = CNNCompressor("facebook/bart-large-cnn", ratio=0.1, min_len=20)
with compressor:
    summaries = compressor.compress(long_docs)
Researchers developed a solid-state lithium battery with a ceramic electrolyte that resists
dendrite growth. Over 1,000 charge cycles the cell retained 92 percent of its capacity. Design
could shorten charging times for electric vehicles.
Parameter Default Meaning
min_len 30 Minimum summary length, in tokens.
num_beams 4 Beam search width.
source_len 512 Fallback input length.

Seq2SeqCompressor

The same as CNNCompressor, with an instruction in front of each document, for general instruction-tuned models such as FLAN-T5.

from topicjev.compress import Seq2SeqCompressor

compressor = Seq2SeqCompressor("google/flan-t5-base", ratio=0.2)
with compressor:
    summaries = compressor.compress(long_docs)
Parameter Default Meaning
instruct "Summarize the following text in two or three plain sentences. {}" Instruction; {} is replaced by the document.
min_len 30 Minimum summary length, in tokens.
num_beams 4 Beam search width.
source_len 1024 Fallback input length.

Note

Small abstractive models can add details that are not in the text. In our test, flan-t5-base summarized the bus-fleet document as being about New York City, which the document never mentions. Check a sample of summaries before building on them, or use LinguaCompressor, which only deletes tokens.

LinguaCompressor

LLMLingua-2 classifies every token as keep or drop and deletes the dropped ones. Nothing is generated, so the result keeps the original wording and cannot invent content. It needs the lingua extra:

pip install -e ".[lingua]"
from topicjev.compress import LinguaCompressor

compressor = LinguaCompressor(
    "microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank",
    ratio=0.5,
)
with compressor:
    pruned = compressor.compress(long_docs)
developed solid - state lithium battery ceramic electrolyte resists dendrite growth.   Over 1, 000
charge cycles cell retained 92 percent capacity,   operated safely temperatures 80 degrees Celsius.
authors design shorten charging times electric vehicles,   manufacturing thin ceramic layers scale expensive.

The output is not fluent prose, but it keeps the original words and every number. Here ratio is the share of tokens to keep. For better quality at a higher cost, use microsoft/llmlingua-2-xlm-roberta-large-meetingbank.

Parameter Default Meaning
target_token -1 Target length in tokens; -1 uses ratio instead.
force_tokens newlines, space, . , : ; ! ? Tokens that are always kept.
force_digits False Always keep tokens that contain digits.
drop_consecutive True Drop repeated force tokens that end up next to each other.

LLMLingua runs on a CUDA GPU when one is available and on the CPU otherwise, including on Apple Silicon Macs.

Compress, then classify

from topicjev.compress import CNNCompressor
from topicjev.entail import XEncoderEntail

with CNNCompressor("facebook/bart-large-cnn", ratio=0.1) as compressor:
    short_docs = compressor.compress(long_docs)

with XEncoderEntail("MoritzLaurer/deberta-v3-base-zeroshot-v2.0") as model:
    results = model.entail(docs=short_docs, lbls=["Battery Technology", "Public Transit", "Sports"])

Each with block frees its model before the next one is loaded, so the two never share GPU memory.