Run your own AI model
Akollo's AI features run on the models Akollo provides until an admin connects something else. This guide is for the case where you want the model on hardware you control: a laptop for a trial, a workstation, or a server in your own network. You install a model runtime, start it as an OpenAI-compatible server, and add it in Akollo as an Own model server connection, directly or through the AI Connector.
Why run your own model
- Data stays with you. Questions and answers go to a machine your organisation runs, not to an outside AI provider.
- You choose the model. You decide which open model runs, in which size, and when it is updated.
- No AI tokens spent. Requests through your own connection do not use your organisation's Akollo AI tokens. See AI access and tokens.
- It is a good first step. A laptop is enough to try it with a small model before you plan a server. For a full installation inside the institution, see Running the model in the institution.
Choose a runtime
Any server that speaks the OpenAI-compatible API works. These three are the most common and are all free to use.
| Runtime | Best for | How you run it | Embeddings |
|---|---|---|---|
llama.cpp (llama-server) | Servers, full control, CPU-only machines | Command line, one model per server process | Yes, as a separate server started with --embeddings |
| Ollama | Quick setup on a laptop or a small server | Background service, models pulled by name | Yes |
| LM Studio | Trying models on a desktop with a graphical app | Desktop app with a built-in local server | Yes |
For a shared server that the whole organisation uses, llama.cpp (or vLLM on a graphics-card server) is usually the better choice. Ollama and LM Studio are the easiest way to try things on one machine.
Install
With Homebrew:
# llama.cpp (includes llama-server)
brew install llama.cpp
# or Ollama
brew install ollamaYou can also download the Ollama app from ollama.com. For LM Studio, download the app from lmstudio.ai and drag it to Applications. Apple silicon Macs use the built-in graphics memory automatically.
Which model for which machine
The models below are the ones Akollo's own sizing is based on. They are open models you can download from Hugging Face or pull by name in Ollama and LM Studio.
| Machine | Model | Use it for |
|---|---|---|
| Laptop with 16 GB of memory | Gemma 4 E4B or Jan-v1-4B | Trying it out, routing, safety checks, short answers |
| Apple silicon with 32–48 GB, or a PC with a 24 GB graphics card | Gemma 4 26B-A4B | Everyday answers and planning for a team or a pilot |
| Any of the above, alongside the chat model | EmbeddingGemma | Search. It is small and runs well on a processor |
Quantisation means storing the model's numbers with fewer bits. A Q4 (4-bit) file is about half the size of a
Q8 (8-bit) file, loads faster and answers faster, with a small loss in quality. Q4 is a good default. Pick Q8 only
if the machine has memory to spare. Model files for llama.cpp have the extension .gguf, and the quantisation is
usually part of the file name (for example Q4_K_M).
| EmbeddingGemma | 0.5 GB |
|---|---|
| Jan-v1-4B | 3 GB |
| Gemma 4 E4B | 4.5 GB |
| Gemma 4 26B-A4B | 16 GB |
For sizing a server for hundreds or thousands of people, see Running the model in the institution.
Start the server
In the examples, <port> is a port you choose and <model-file> is the model you downloaded. The OpenAI-compatible
endpoint is the server address followed by /v1.
Start one server for answers and, if you want to use your own model for search, a second one for embeddings:
# Chat model
llama-server -m <model-file>.gguf --port <port>
# Or let llama-server download a model from Hugging Face
llama-server -hf <publisher>/<model>-GGUF:Q4_K_M --port <port>
# Embedding model, on its own port
llama-server -m <embedding-model-file>.gguf --embeddings --port <port>To protect the server with a key, add --api-key <your-key>. Enter the same key in Akollo.
Who can reach the server
By default, all three runtimes listen only on the machine itself (localhost). That is the safest setting and is enough when the AI Connector runs on the same machine.
| You want | What to do |
|---|---|
| Only this machine | Keep the default. |
| Other machines in your network | llama.cpp: add --host 0.0.0.0. Ollama: set the OLLAMA_HOST environment variable to 0.0.0.0. LM Studio: turn on serving on the local network in the server settings. |
| Akollo connects directly | Put the server behind HTTPS with a valid certificate (for example a reverse proxy), protect it with a key, and ask your installation's operator to add its address to the allowed destinations. |
Warning
Never expose a model server to the internet without a key. Anyone who can reach it can use your hardware. With
llama.cpp use --api-key. For Ollama and LM Studio, use a reverse proxy that checks a key, or keep the server
private and use the AI Connector.
Connect it in Akollo
Organisation owners and admins manage connections. See AI connections for the whole page.
Open the connections page
Open Settings › AI › Connections.
Add the server
Under Add a connection, choose Own model server as the Provider and give it a Name. Under Address,
enter the server's HTTPS address including /v1, for example https://<your-model-server>/v1. If you set a key on
the server, paste it under API key. Otherwise leave it empty.
Choose which data may go through it
Under Data allowed, pick Public data only, Up to internal data or Up to confidential data, then choose Add connection.
Find models and test
Choose Find models. The list shows the models your server serves. Pick a model and choose Test. The test sends a short set of fixed requests: answers in Turkish, structured answers and using a tool. The connection can be activated only after it passes.
Activate
Choose Activate. In live organisations, a second admin approves the activation on the same page.
Choose which jobs use it
Under Which model does which job, choose your connection and model for the jobs you want, then choose Save. Jobs you leave alone keep using the models Akollo provides.
A sensible split with the models above:
| Job | Suggested model |
|---|---|
| Routing | Jan-v1-4B |
| Safety check | Gemma 4 E4B |
| Everyday answers | Gemma 4 26B-A4B |
| Planning and complex work | Gemma 4 26B-A4B, or keep the models Akollo provides |
| Search | EmbeddingGemma |
Note
A small model may fail the tool or structured-answer checks. Use it for Routing or Safety check and give Everyday answers a larger model.
Servers inside a company network
From a cloud installation, Akollo does not connect to private network addresses. If your model server sits inside your company network, or you do not want to open an inbound port, use the AI Connector. It is a small container that runs inside your network next to the model server. It connects out to Akollo over HTTPS, picks up requests, sends them to the model server and returns the answers. No inbound port is opened.
Add the connector in Akollo
On Settings › AI › Connections, under Add an AI Connector, give it a name and choose Add connector.
Copy the token
Copy the Connector token. It is shown only once. Keep it in your secret store.
Start the container
Run the AI Connector container next to the model server. It needs three settings: Akollo's address, the token, and the model server's OpenAI-compatible address. A fourth, optional setting holds the model server's key if it needs one.
docker run --restart unless-stopped \
-e AKOLLO_URL=https://<your-akollo-address> \
-e AKOLLO_CONNECTOR_TOKEN=<connector-token> \
-e AKOLLO_CONNECTOR_TARGET=http://<model-server>:<port>/v1 \
-e AKOLLO_CONNECTOR_TARGET_KEY=<model-server-key> \
<ai-connector-image>Check, test and activate
The connection card shows Connector online once it connects. Then use Find models, Test and Activate as for a direct server.
- The connector only needs outgoing HTTPS to Akollo and a route to the model server. The model server itself can stay on plain HTTP inside your network.
- Run one copy per token, restarted automatically if it stops. A second copy with the same token is refused.
- The connector forwards model listing and chat answers. For the Search job, use a direct Own model server connection or the models Akollo provides.
- New token stops the old token at once. Restart the connector with the new token, then test and activate the connection again.
- The connector logs one line per request and never the token or the content.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The test says the connection could not be reached | Wrong address, missing /v1, wrong key, or the server is not running | Open the address followed by /models from a machine that should reach it. Check the key, then test again. |
| The address is not accepted | Not HTTPS, not on the installation's allowed list, or a private network address | Put the server behind HTTPS and ask your operator to allow the address, or use the AI Connector. |
| Find models shows no models | No model loaded, or the key belongs to another server | Load or pull a model (in LM Studio, load it in the server view), then choose Find models again. |
| The test fails some checks | The model is too small for structured answers or tool use | Use a larger model, such as Gemma 4 26B-A4B, for Everyday answers. Keep small models for Routing and Safety check. |
| Answers are slow or time out | The model is too large for the machine, or it runs only on the processor | Use a smaller model or a Q4 file, make sure the graphics card is used, and close other heavy programs. |
| The server stops with an out-of-memory error | Model plus context does not fit in memory | Use a smaller model or quantisation, or shorten the context window. |
| Search results are poor or fail | A chat model was chosen for Search, or the embedding server is not running | Choose an embedding model such as EmbeddingGemma for Search. With llama.cpp, start it with --embeddings. Do not switch embedding models back and forth. |
| Certificate or TLS errors | Self-signed or expired certificate | Use a certificate your installation trusts, or use the AI Connector so Akollo does not connect to the server directly. |
| Browser or CORS errors when testing by hand | The runtime blocks requests from web pages | This does not affect Akollo, which calls the server from its own servers. Test with a command-line tool instead. |
| Other machines cannot reach the server | It listens only on localhost, or a firewall blocks the port | Bind it to the network (see above) and open the port only to the machines that need it. |
| Connector offline | Wrong token, Akollo's address not reachable, or the token was replaced | Check the container log, the three settings and outgoing HTTPS. After New token, restart with the new token. |
FAQ
No. Small models such as Jan-v1-4B, Gemma 4 E4B and EmbeddingGemma run on a processor. A graphics card, or an Apple silicon Mac with enough memory, makes the larger Gemma 4 26B-A4B fast enough for everyday answers.
No. Requests through your own connection do not use your organisation's Akollo AI tokens.
Yes. Install Ollama or LM Studio, load a small model, and connect it through the AI Connector running on the same laptop. Use a model your laptop can hold, such as Gemma 4 E4B.
Not with the AI Connector. It connects out to Akollo, so no inbound port is opened. A direct Own model server connection needs an HTTPS address that Akollo can reach.
Yes. Under Which model does which job, choose a model for each job separately. Jobs you leave alone keep using the models Akollo provides.
Requests that use it fail until it is back. To move jobs back to the models Akollo provides, change them under Which model does which job. Disable turns the connection off for good.