Local inference and operations

Keep the app and model on your machine.

IdeaForge can use Ollama, LM Studio, or another OpenAI-compatible server on loopback. The simplest and most reliable arrangement serves the static app locally too, so the browser never has to cross from a public HTTPS page into a local HTTP service.

The browser talks to the model. The IdeaForge container is only an nginx static-file server. It never receives an API key, transcript, model request, or model response.

Choose a topology

Topology Use it when Important consequence
Local static server + native model server You already run Ollama or LM Studio on the host. Easiest browser path. Both page and model are local, so Chrome's public-page local-network permission is not involved.
Source Compose stack You have a clone and want the app image built from the working tree. Runs nginx and Ollama, publishes both on 127.0.0.1, and keeps models in a named volume.
Release Compose stack You want published images without cloning or building. Uses the released IdeaForge image and Ollama image. The supplied GPU block targets NVIDIA container support.
App-only container + native model server Ollama already uses the host GPU, especially on Apple Silicon or a setup not covered by the repository's NVIDIA Compose block. Containerises only the static app. The browser still reaches the native model at a host loopback address.
Hosted app + local model You cannot or do not want to serve the app locally. Requires server CORS plus Chrome local-network permission. The shipped app treats this path as unsupported in Safari; serve the app locally instead.

Run from a clone without Docker

There are no dependencies and no build step. Node is used for the repository tools; the app itself only needs a static server over a real origin.

npm run serve
# open http://127.0.0.1:8765

Do not open index.html with file://. ES modules, IndexedDB, installability, and the service worker all depend on an origin.

  1. Start Ollama, LM Studio, or another OpenAI-compatible server.
  2. Open http://127.0.0.1:8765.
  3. Select the matching local provider.
  4. Prefer a server address such as http://127.0.0.1:11434/v1. Press Check the connection.
  5. Choose an installed model from the suggestions, or type its exact identifier. Leave the wrap-up field blank and the same model runs that call too.

Run the source Compose stack

docker compose -f docker/compose.yml up -d
docker compose -f docker/compose.yml exec ollama ollama pull <model-name>
# open http://127.0.0.1:8765

The app is built from the current tree with docker/Dockerfile.app. The image copies only the publishable static tree into nginx. Ollama's model store is a named volume, so an ordinary docker compose down does not discard downloaded model data. Do not add -v unless removing that volume is intentional.

If port 11434 is already occupied, publish the container on a different host port and put that same address in IdeaForge:

IDEAFORGE_OLLAMA_PORT=11435 docker compose -f docker/compose.yml up -d
# Server address: http://127.0.0.1:11435/v1

In PowerShell, set the variable first:

$env:IDEAFORGE_OLLAMA_PORT = '11435'
docker compose -f docker/compose.yml up -d

Run published images

The standalone release Compose file needs no clone and no local app build:

curl -fsSLO https://raw.githubusercontent.com/jaypetez/ideaforge/main/docker/compose.release.yml
docker compose -f compose.release.yml up -d
docker compose -f compose.release.yml exec ollama ollama pull <model-name>
# open http://127.0.0.1:8765

To refresh an existing installation, pull before recreating the services:

docker compose -f compose.release.yml pull
docker compose -f compose.release.yml up -d

The release file shares the same Compose project and model-volume name as the source stack. This avoids downloading the model again when switching between a released app and a locally built one, but it also means the two stacks should not run at the same time.

Containerise only the static app

docker run -d --name ideaforge \
  -p 127.0.0.1:8765:80 \
  ghcr.io/jaypetez/ideaforge:latest

Run Ollama or LM Studio natively, then point IdeaForge at its loopback API. This is the practical route when the host's native model runtime has the right GPU support and a Linux container does not, including the normal Apple Silicon arrangement.

GPU and CPU behaviour

Both repository Compose definitions reserve NVIDIA GPU devices. Docker does not silently ignore a missing NVIDIA device driver: service creation fails with a clear device-driver error. That is intentional, because an unnoticed CPU fallback looks like an extremely slow application rather than a deployment fault.

See Ollama's Docker guidance and Docker's Compose GPU guidance for host prerequisites.

Use an address IdeaForge can safely reach

Address Result Reason
http://127.0.0.1:<port>/v1 Recommended Explicit IPv4 loopback avoids a host name resolving to a different local server in Node and the browser.
http://localhost:<port>/v1 Accepted, but can be ambiguous On a machine with native and containerised servers, IPv4 and IPv6 resolution can select different processes.
http://[::1]:<port>/v1 Refused by the UI The address is loopback, but the page's CSP cannot express an IPv6 literal host source reliably. Use 127.0.0.1.
http://0.0.0.0:<port>/v1 Refused 0.0.0.0 is a bind address, not a loopback destination.
http://192.168.x.x:<port>/v1 Refused Arbitrary LAN hosts are outside the app's key-exfiltration boundary.
http://localhost.evil.example/v1 Refused IdeaForge parses the URL hostname; a string that merely starts with localhost is not trusted.
On a phone, localhost is the phone. It never means a laptop on the same Wi-Fi. IdeaForge deliberately refuses the laptop's LAN address. Use a hosted provider on the phone, or run the interview on the same computer as the local model.

CORS, local-network permission, and Safari

When the app is local

A page at http://127.0.0.1:8765 calling a loopback model stays within the local address space. Ollama's normal localhost origins cover this route. LM Studio still needs CORS enabled for browser access.

When the app is hosted over HTTPS

Two independent gates must pass:

  1. The model server must allow the page's origin. For Ollama, set OLLAMA_ORIGINS to the exact origin that hosts IdeaForge and restart Ollama. The origin has a scheme and host, but no path; for the public Pages deployment it is https://jaypetez.github.io. Ollama reads the environment when it starts, so changing another shell without restarting the running service has no effect.
  2. Chrome must grant Local Network Access. IdeaForge marks the fetch with targetAddressSpace: "local". Chrome shows its permission prompt only after you explicitly press Check the connection; allow it for the site.

LM Studio exposes CORS in its server settings and through the lms server start options. Ollama documents allowed origins in its official FAQ. Chrome documents the current permission model in Local Network Access.

Use a locally served app with Safari. The shipped hosted-HTTPS to local-HTTP path is not supported there. Safari does not provide the Chrome permission route that IdeaForge uses to make this request. Moving the app to http://127.0.0.1:8765 keeps both ends local and avoids that mixed-content boundary. See the general mixed-content model and the long-running WebKit loopback issue.

Validate a real model, not a graceful fallback

IdeaForge is designed to continue from its built-in question bank when inference fails. That resilience makes a superficial end-to-end check unsafe: an interview can finish and export even if no model request succeeded. The repository's validate:local harness therefore checks both application state and what actually left the browser.

Run against a native or already-running Ollama

In a POSIX shell:

OLLAMA_URL=http://127.0.0.1:11434 \
IDEAFORGE_MODEL=<model-name> \
npm run validate:local

In PowerShell:

$env:OLLAMA_URL = 'http://127.0.0.1:11434'
$env:IDEAFORGE_MODEL = '<model-name>'
npm run validate:local

Validate a running app container

Set IDEAFORGE_URL so the harness drives that origin instead of serving the working tree:

$env:IDEAFORGE_URL = 'http://127.0.0.1:8765'
$env:OLLAMA_URL = 'http://127.0.0.1:11434'
$env:IDEAFORGE_MODEL = '<model-name>'
npm run validate:local

Run the containerised validator

docker compose -f docker/compose.yml -f docker/compose.validate.yml run --rm validate

The validation container shares Ollama's network namespace. That makes 127.0.0.1:11434 genuinely point to Ollama inside the container while preserving IdeaForge's loopback-only policy. It does not create a special test-only remote-server exception.

What the validator proves

Useful validation controls

Variable Purpose
OLLAMA_URL Names the exact Ollama server to inspect and use.
IDEAFORGE_MODEL Selects the already-pulled model to validate.
IDEAFORGE_URL Drives an already-served app or container instead of the working tree.
VALIDATE_TURNS Changes the interview length before the requested wrap-up.
VALIDATE_TURN_MS Changes the per-turn allowance for unusually slow local hardware.
VALIDATE_MIN_GBPS Overrides the decode memory-bandwidth floor.
VALIDATE_ALLOW_CPU=1 Continues the application checks after an intentional CPU-only result. It does not pretend that CPU execution passed the GPU claim.
VALIDATE_MODE=handsfree Drives the same real model through scripted speech input and output. Run typed and hands-free as separate commands.
CHROME_PATH Points the harness at Chrome or Chromium when automatic discovery fails.

For the containerised run, pass controls that are not declared in the Compose file with docker compose run -e, for example -e VALIDATE_MODE=handsfree or -e VALIDATE_ALLOW_CPU=1.

Implementation references

The development stack is docker/compose.yml, the standalone release stack is docker/compose.release.yml, and the validation overlay is docker/compose.validate.yml. The real-model checks themselves are in tools/validate-local.mjs.

Loopback parsing and Chrome's fetch option are in src/providers/http.js, while local presets, model discovery, CORS diagnostics, and Safari guidance are in src/providers/openaiCompat.js.