topicjev.gen¶
Embeddings and structured text generation. See the Embeddings and LLM Generation guides.
Embeddings¶
Embedder
¶
Bases: ModelBackend
Base class for text embedding backends.
Source code in topicjev/gen/embed.py
embed_docs
abstractmethod
¶
get_embeds
¶
Compute embeddings for a list of documents.
Source code in topicjev/gen/embed.py
LocalEmbedder
¶
Bases: Embedder
Embeds text using local SentenceTransformer models.
Source code in topicjev/gen/embed.py
load_model
¶
Load SentenceTransformer model on detected device.
Source code in topicjev/gen/embed.py
embed_docs
¶
Compute document embeddings.
Source code in topicjev/gen/embed.py
close
¶
Offload embedding model to CPU and clear GPU memory.
Source code in topicjev/gen/embed.py
Generation¶
Generator
¶
Generator(
mname: str,
*,
temperature: float = 0.1,
max_tokens: int = 4000,
max_input_tokens: int = 4000,
json_mode: bool = True,
reason: bool = False,
batch_size: int = 4,
max_retries: int = 5,
**kwargs: Any,
)
Bases: ModelBackend
Base class for prompt-to-text generation and structured JSON extraction.
Source code in topicjev/gen/chat.py
query
abstractmethod
¶
query(
user_prompt: str,
sys_prompt: str = "You are a helpful AI assistant",
) -> Union[str, dict[str, Any]]
Query the model with a single prompt and optional system instructions.
batch
abstractmethod
¶
Execute a prompt template over a batch of input documents.
run_chain
¶
Alias for batch() for pipeline compatibility.
LocalGenerator
¶
Bases: CausalLMBackend, Generator
Generates text using a local causal language model on hardware devices.
Source code in topicjev/gen/chat.py
tok
property
writable
¶
Alias for tokenizer for consistency across local backends.
load_model
¶
Load causal LM, left-padded tokenizer, and optionally compile for CUDA.
Source code in topicjev/gen/chat.py
batch_gen
¶
Generate text outputs for a list of formatted prompts.
Source code in topicjev/gen/chat.py
query
¶
query(
user_prompt: str,
sys_prompt: str = "You are a helpful AI assistant",
) -> Union[str, dict[str, Any]]
Run single-prompt inference.
Source code in topicjev/gen/chat.py
batch
¶
Batch-generate responses across input documents.
Source code in topicjev/gen/chat.py
close
¶
Offload local model and clear GPU cache.
Source code in topicjev/gen/chat.py
clean_json
¶
Extract and parse a JSON dictionary from an LLM response string.