Setting Up local LLM with Ollama

Useful links and resources:

Download and install Ollama

There are 2 modes of installation available. You can download the installer for Mac, Linux or Windows from their website or your can run this in your Terminal on Mac or Linux:

curl -fsSL https://ollama.com/install.sh | sh

Or PowerShell on Windows:

irm https://ollama.com/install.ps1 | iex

This will install the necessary components, but loading a model is a separate process. First we'd need to figure out the computer's specs to make sure the model will be able to run on your machine.

Find your system resource information

Both Windows and Mac machines can run local AI models, but they do it differently and therefore have different system requirements. Apple introduced shared memory architecture for their M-series processors (M1, M2, etc.) This means that the memory can be allocated between the central processor and the additional graphics and neural engines as needed. This makes it a very powerful and flexible architecture, and it is usually preferred for running larger models. Windows machines have separate memory for the CPU (central processor) and the GPU (graphics processor that also runs AI models). GPUs from NVIDIA have the ability to run their custom CUDA architecture, which is often a preferred framework for running AI models. This means that even though the amount of memory available to the GPU might be smaller than on a Mac, the software architecture is very fine-tuned and that smaller models can run faster.

Choose you initial model

We'll start by setting up a simple chat model just to make sure your set-up is operational.