topicjev.backend¶
Shared device handling and model lifecycle. Every backend in TopicJev derives from
ModelBackend. See Hardware and Memory.
Devices¶
detect_device
¶
Select the optimal torch device and dtype for local inference.
Prefers CUDA, then Apple Silicon (MPS), then CPU. Set the TOPICJEV_DEVICE
environment variable (e.g. cpu, mps, cuda:1) to override.
Source code in topicjev/backend.py
empty_device_cache
¶
Release cached accelerator memory (CUDA or MPS) after a model is offloaded.
Lifecycle¶
ModelBackend
¶
Bases: ABC
Base contract for model and API client lifecycles.
Source code in topicjev/backend.py
load_model
abstractmethod
¶
LocalModelBackend
¶
Bases: ModelBackend
Base for local PyTorch-based models with device and memory management.
Source code in topicjev/backend.py
close
¶
Offload model to CPU and clear GPU cache if allocated.
Source code in topicjev/backend.py
CausalLMBackend
¶
Bases: LocalModelBackend, ChatTemplateMixin
Base for local causal language models with left-padded tokenizers.
Source code in topicjev/backend.py
load_model
¶
Load causal LM and left-padded tokenizer onto target hardware device.
Source code in topicjev/backend.py
Prompt formatting¶
ChatTemplateMixin
¶
Mixin providing chat-template prompt formatting for causal / chat models.
format_chat_prompt
¶
format_chat_prompt(
tokenizer: Any,
text: str,
sys_prompt: str | None = None,
*,
chat: bool = True,
thinking: bool = False,
add_generation_prompt: bool = True,
) -> str
Format input text using the tokenizer's chat template if available.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tokenizer
|
Any
|
Hugging Face tokenizer instance. |
required |
text
|
str
|
User query or prompt text. |
required |
sys_prompt
|
str | None
|
Optional system prompt to prepend. |
None
|
chat
|
bool
|
Whether to apply chat templating (if supported by tokenizer). |
True
|
thinking
|
bool
|
Whether to keep thinking enabled for reasoning-tuned models. |
False
|
add_generation_prompt
|
bool
|
Whether to append the generation prompt. |
True
|
Returns:
| Type | Description |
|---|---|
str
|
Formatted prompt string. |
Source code in topicjev/backend.py
resolve_max_context_length
¶
Determine the effective maximum sequence length from tokenizer or model config.