5. Ollama Local AI
Know what Ollama is and how TPEE should describe it safely.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
Ollama is a separate local model runner. It can run models such as Llama, Qwen, Mistral, and other open-weight models on your own PC.
TPEE does not start Ollama, call its HTTP API, download models, or manage the Ollama service. TPEE only helps you write prompts and notes that you may paste into your own Ollama workflow.
Local AI quality depends heavily on RAM, VRAM, quantisation, and model size. Smaller quantised models are faster and cheaper, but may reason less deeply.
A financial analyst at a brokerage runs a local model to summarise internal market notes so client information stays on the office machines, then pastes the finished TPEE prompt in by hand.
A security operations analyst drafts incident prompts in TPEE and copies them into a locally run model, keeping vulnerability details away from any external service.
A clinic practice manager writes recall-reminder prompts in TPEE and moves them manually into a local model that runs on the reception computer.
A secondary school teacher writes a feedback prompt in TPEE, then pastes it into a local model on the classroom laptop so learner names never leave the school network.
A lawyer uses a local model for first-pass summaries of confidential correspondence, copying each finished prompt from TPEE into the local runtime by hand.
A manufacturing quality engineer drafts inspection-report prompts in TPEE and runs them on a local model because the drawings carry restricted information.
A municipal records clerk keeps a local model for summarising internal memos, writing the prompt in TPEE and pasting it across since TPEE never contacts the model itself.
A veterinary surgeon drafts anaesthetic-note summaries in TPEE and copies them into a local model running on the clinic workstation.
A journalist drafts first-pass interview summaries in TPEE, then pastes the prompt into a local model so unpublished source notes stay on the newsroom machine.
A robotics technician writes maintenance prompts in TPEE and copies them into a local model while the workshop network stays disconnected.
Check yourself
Question 1: What is Ollama in this lesson?
- A TPEE internal server
- A cloud billing provider
- A Rust compiler
- A separate local model runner — correct
Answer: A separate local model runner
Ollama runs separately from TPEE and can host local models.
Question 2: What does TPEE do with Ollama?
- Downloads models automatically
- Calls the Ollama HTTP API
- Prepares prompts and notes for manual use — correct
- Starts the Ollama service
Answer: Prepares prompts and notes for manual use
TPEE must not start or call Ollama; the user runs it separately if desired.
Question 3: What usually limits larger Ollama models first?
- Printer ink
- RAM or VRAM capacity — correct
- Mouse speed
- Screen color
Answer: RAM or VRAM capacity
Local model size and KV cache must fit available memory.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.