Skip to content

Tag: ollama

All the articles with the tag "ollama".

Ollama Keep Alive: Unload Models from VRAM

Ollama Keep Alive: Unload Models from VRAM

· Updated:

Ollama holds models in VRAM after every request. Set keep_alive, force an unload through the API, and cap how many models stay resident at once.

Key Parameters of Large Language Models

Key Parameters of Large Language Models

· Updated:

Temperature, top-p, top-k, context length, LLM inference parameters explained so you stop guessing why the model gives weird output.

Exploring the Diverse World of LLM Models

Exploring the Diverse World of LLM Models

· Updated:

LLaMA, Mistral, Falcon, GPT, the LLM landscape is crowded. Compare model families, sizes, licensing, and what each is actually good for.