Local inference and operations
Keep the app and model on your machine.
IdeaForge can use Ollama, LM Studio, or another OpenAI-compatible server on loopback. The simplest and most reliable arrangement serves the static app locally too, so the browser never has to cross from a public HTTPS page into a local HTTP service.
Choose a topology
| Topology | Use it when | Important consequence |
|---|---|---|
| Local static server + native model server | You already run Ollama or LM Studio on the host. | Easiest browser path. Both page and model are local, so Chrome's public-page local-network permission is not involved. |
| Source Compose stack | You have a clone and want the app image built from the working tree. |
Runs nginx and Ollama, publishes both on 127.0.0.1, and keeps
models in a named volume.
|
| Release Compose stack | You want published images without cloning or building. | Uses the released IdeaForge image and Ollama image. The supplied GPU block targets NVIDIA container support. |
| App-only container + native model server | Ollama already uses the host GPU, especially on Apple Silicon or a setup not covered by the repository's NVIDIA Compose block. | Containerises only the static app. The browser still reaches the native model at a host loopback address. |
| Hosted app + local model | You cannot or do not want to serve the app locally. | Requires server CORS plus Chrome local-network permission. The shipped app treats this path as unsupported in Safari; serve the app locally instead. |
Run from a clone without Docker
There are no dependencies and no build step. Node is used for the repository tools; the app itself only needs a static server over a real origin.
npm run serve
# open http://127.0.0.1:8765
Do not open index.html with file://. ES modules,
IndexedDB, installability, and the service worker all depend on an origin.
- Start Ollama, LM Studio, or another OpenAI-compatible server.
- Open
http://127.0.0.1:8765. - Select the matching local provider.
-
Prefer a server address such as
http://127.0.0.1:11434/v1. Press Check the connection. - Choose an installed model from the suggestions, or type its exact identifier. Leave the wrap-up field blank and the same model runs that call too.
Run the source Compose stack
docker compose -f docker/compose.yml up -d
docker compose -f docker/compose.yml exec ollama ollama pull <model-name>
# open http://127.0.0.1:8765
The app is built from the current tree with
docker/Dockerfile.app.
The image copies only the publishable static tree into nginx. Ollama's model store
is a named volume, so an ordinary docker compose down does not discard
downloaded model data. Do not add -v unless removing that volume is
intentional.
If port 11434 is already occupied, publish the container on a different
host port and put that same address in IdeaForge:
IDEAFORGE_OLLAMA_PORT=11435 docker compose -f docker/compose.yml up -d
# Server address: http://127.0.0.1:11435/v1
In PowerShell, set the variable first:
$env:IDEAFORGE_OLLAMA_PORT = '11435'
docker compose -f docker/compose.yml up -d
Run published images
The standalone release Compose file needs no clone and no local app build:
curl -fsSLO https://raw.githubusercontent.com/jaypetez/ideaforge/main/docker/compose.release.yml
docker compose -f compose.release.yml up -d
docker compose -f compose.release.yml exec ollama ollama pull <model-name>
# open http://127.0.0.1:8765
To refresh an existing installation, pull before recreating the services:
docker compose -f compose.release.yml pull
docker compose -f compose.release.yml up -d
The release file shares the same Compose project and model-volume name as the source stack. This avoids downloading the model again when switching between a released app and a locally built one, but it also means the two stacks should not run at the same time.
Containerise only the static app
docker run -d --name ideaforge \
-p 127.0.0.1:8765:80 \
ghcr.io/jaypetez/ideaforge:latest
Run Ollama or LM Studio natively, then point IdeaForge at its loopback API. This is the practical route when the host's native model runtime has the right GPU support and a Linux container does not, including the normal Apple Silicon arrangement.
GPU and CPU behaviour
Both repository Compose definitions reserve NVIDIA GPU devices. Docker does not silently ignore a missing NVIDIA device driver: service creation fails with a clear device-driver error. That is intentional, because an unnoticed CPU fallback looks like an extremely slow application rather than a deployment fault.
- For the release file, follow the marked GPU BLOCK instructions in the file if you deliberately want CPU-only Ollama.
- For other GPU stacks, prefer native Ollama plus the app-only container, or maintain a local Compose override appropriate to that host. The repository file does not claim to configure every GPU runtime.
- A CPU model can still answer correctly, but cold loads and long prompts may take much longer. Use a model that fits available memory rather than treating repeated request deadlines as a browser problem.
See Ollama's Docker guidance and Docker's Compose GPU guidance for host prerequisites.
Use an address IdeaForge can safely reach
| Address | Result | Reason |
|---|---|---|
http://127.0.0.1:<port>/v1 |
Recommended | Explicit IPv4 loopback avoids a host name resolving to a different local server in Node and the browser. |
http://localhost:<port>/v1 |
Accepted, but can be ambiguous | On a machine with native and containerised servers, IPv4 and IPv6 resolution can select different processes. |
http://[::1]:<port>/v1 |
Refused by the UI |
The address is loopback, but the page's CSP cannot express an IPv6 literal
host source reliably. Use 127.0.0.1.
|
http://0.0.0.0:<port>/v1 |
Refused |
0.0.0.0 is a bind address, not a loopback destination.
|
http://192.168.x.x:<port>/v1 |
Refused | Arbitrary LAN hosts are outside the app's key-exfiltration boundary. |
http://localhost.evil.example/v1 |
Refused |
IdeaForge parses the URL hostname; a string that merely starts with
localhost is not trusted.
|
CORS, local-network permission, and Safari
When the app is local
A page at http://127.0.0.1:8765 calling a loopback model stays within
the local address space. Ollama's normal localhost origins cover this route. LM
Studio still needs CORS enabled for browser access.
When the app is hosted over HTTPS
Two independent gates must pass:
-
The model server must allow the page's origin.
For Ollama, set
OLLAMA_ORIGINSto the exact origin that hosts IdeaForge and restart Ollama. The origin has a scheme and host, but no path; for the public Pages deployment it ishttps://jaypetez.github.io. Ollama reads the environment when it starts, so changing another shell without restarting the running service has no effect. -
Chrome must grant Local Network Access.
IdeaForge marks the fetch with
targetAddressSpace: "local". Chrome shows its permission prompt only after you explicitly press Check the connection; allow it for the site.
LM Studio exposes CORS in its server settings and through the
lms server start options.
Ollama documents allowed origins in its
official FAQ. Chrome documents the current
permission model in
Local Network Access.
http://127.0.0.1:8765 keeps both ends local and avoids that
mixed-content boundary. See the general
mixed-content model
and the long-running
WebKit loopback issue.
Validate a real model, not a graceful fallback
IdeaForge is designed to continue from its built-in question bank when inference
fails. That resilience makes a superficial end-to-end check unsafe: an interview can
finish and export even if no model request succeeded. The repository's
validate:local harness therefore checks both application state and what
actually left the browser.
Run against a native or already-running Ollama
In a POSIX shell:
OLLAMA_URL=http://127.0.0.1:11434 \
IDEAFORGE_MODEL=<model-name> \
npm run validate:local
In PowerShell:
$env:OLLAMA_URL = 'http://127.0.0.1:11434'
$env:IDEAFORGE_MODEL = '<model-name>'
npm run validate:local
Validate a running app container
Set IDEAFORGE_URL so the harness drives that origin instead of serving
the working tree:
$env:IDEAFORGE_URL = 'http://127.0.0.1:8765'
$env:OLLAMA_URL = 'http://127.0.0.1:11434'
$env:IDEAFORGE_MODEL = '<model-name>'
npm run validate:local
Run the containerised validator
docker compose -f docker/compose.yml -f docker/compose.validate.yml run --rm validate
The validation container shares Ollama's network namespace. That makes
127.0.0.1:11434 genuinely point to Ollama inside the container while
preserving IdeaForge's loopback-only policy. It does not create a special
test-only remote-server exception.
What the validator proves
- Ollama answers, and the requested model is actually pulled.
- A host name does not resolve to two different Ollama servers with different model fingerprints.
- The model loads before the browser run, and its reported model size is resident in GPU memory when GPU validation is required.
- A warm decode sample is fast enough for the configured memory-bandwidth floor, without misclassifying a two-token first response as a speed test.
- The real app boots with no uncaught error or CSP violation.
- The browser reaches the same Ollama instance as the Node harness.
- The app reads the real model list and sees the selected model.
-
Every model-sourced question and the final synthesis produce actual browser POST
requests; saved
questionSourcevalues agree. - No turn silently falls back to the checklist, the refined prompt is generated, and the model remains in VRAM through the wrap-up when still loaded.
Useful validation controls
| Variable | Purpose |
|---|---|
OLLAMA_URL |
Names the exact Ollama server to inspect and use. |
IDEAFORGE_MODEL |
Selects the already-pulled model to validate. |
IDEAFORGE_URL |
Drives an already-served app or container instead of the working tree. |
VALIDATE_TURNS |
Changes the interview length before the requested wrap-up. |
VALIDATE_TURN_MS |
Changes the per-turn allowance for unusually slow local hardware. |
VALIDATE_MIN_GBPS |
Overrides the decode memory-bandwidth floor. |
VALIDATE_ALLOW_CPU=1 |
Continues the application checks after an intentional CPU-only result. It does not pretend that CPU execution passed the GPU claim. |
VALIDATE_MODE=handsfree |
Drives the same real model through scripted speech input and output. Run typed and hands-free as separate commands. |
CHROME_PATH |
Points the harness at Chrome or Chromium when automatic discovery fails. |
For the containerised run, pass controls that are not declared in the Compose file
with docker compose run -e, for example
-e VALIDATE_MODE=handsfree or
-e VALIDATE_ALLOW_CPU=1.
Implementation references
The development stack is
docker/compose.yml,
the standalone release stack is
docker/compose.release.yml,
and the validation overlay is
docker/compose.validate.yml.
The real-model checks themselves are in
tools/validate-local.mjs.
Loopback parsing and Chrome's fetch option are in
src/providers/http.js,
while local presets, model discovery, CORS diagnostics, and Safari guidance are in
src/providers/openaiCompat.js.