Changelog¶
TopicJev is in early development: versions 0.x may change the API, and .devN versions are test
releases.
0.1.0.dev3¶
2026-10-05
- Repository moved to anatems1/TopicJev.
- New examples in
examples/, both on a 1,500-article slice of AG News loaded with Hugging Facedatasets(replacingbertopic-example.pyand the bundled PNAS data):bertopic_laya.py: BERTopic topics, verified withLayaEntail;kmeans_genlbs_laya.py: KMeans clusters, labeled byQwen3-1.7Bthrough thereduce→classesprompts, then verified withLayaEntail.
- The
examplesextra addsdatasets. - README rewritten: motivation, use cases, a road-feedback Quick Start, the examples, current limitations and acknowledgements; new logo.
- Docstrings cite LLMLingua-2 (
LinguaCompressor) and GoalEx (GoalExEntail).
0.1.0.dev2¶
2026-10-02
- NVIDIA GPUs on Windows and Linux:
requirements-cuda.txtinstalls PyTorch 2.14.1 built for CUDA 12.6 from the PyTorch package index, then TopicJev with all extras. The default Windows PyTorch from PyPI is CPU-only, so an NVIDIA GPU sat idle. The README explains how to check the driver withnvidia-smi. - Package metadata: SPDX license
MITwithlicense-files, authors, keywords and classifiers; requires hatchling 1.27 or newer. - The version is read from
topicjev/__init__.pyonly. ruffmoved from a publisheddevextra to adevdependency group.- Ships
py.typed, so type checkers use TopicJev's type hints.
0.1.0.dev1¶
2026-10-02
- Fixed: the wheel did not build (
python -m buildandpip install .failed; editable installs were unaffected). The bundled prompts were included twice. - The BERTopic example works on fresh clones, Apple Silicon and transformers 5:
nomic-embed-text-v1.5is loaded through transformers' built-in NomicBert, because the checkpoint's own code fails on transformers 5.LocalEmbeddergained atrust_remote_codeoption (defaultTrue).- Settings are constants; results go to
output/<config hash>/with aconfig.json, and the folder is created if missing. UMAP and KMeans are seeded fromRND_SEED. Embeddings use nomic'sclustering:prefix.
- Devices:
detect_device()prefers CUDA, then MPS, then the CPU, andTOPICJEV_DEVICEoverrides it. MPS runs infloat32. empty_device_cache()frees the CUDA or MPS cache; everyclose()uses it.- Laya and LLMLingua keep the CUDA device index (e.g.
cuda:1). - Requires transformers 5.5 or newer and sentence-transformers 6.0 or newer.
Initial release¶
2026-10-01
topicjev.compress,topicjev.entailandtopicjev.genpipelines.