Skip to content
HN On Hacker News ↗

Ollaya — Run decision models locally.

▲ 618 points • 145 comments • by Ardakilic • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

77 %

AI likelihood · overall

Mixed
23% human-written 77% AI-generated
SEGMENTS · HUMAN 0 of 3
SEGMENTS · AI 1 of 3
WORD COUNT 498
PEAK AI % 87% · §1
Analyzed
Sep 25
backend: pangram/v3.3
Segments scanned
3 windows
avg 166 words each
Distribution
23 / 77%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 498 words · 3 segments analyzed

Human AI-generated
§1 AI · 87%

Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware.DownloadBrowse models ollaya run laya --preset triage \"I was charged twice this month and want a refund."Answers returned by the modelQuestionAnswerProbabilityintentrefund1.00is_urgentno0.87frustration1.59 / 3 clearly annoyed0.36refund_requestedyes0.88churn_riskno0.89Real output: routed to laya:en, answered in 8.9 ms on an RTX 4090.FastDrop-in compatibleOpen modelsYour data stays yoursPlatformsFastDecisions in milliseconds.A decision model answers in a single forward pass, with no token-by-token generation. On your own GPU, a five-question request to Laya takes about 10 ms, end to end through the HTTP API.Every model, one scale · median latency, lower is betterlaya:multilingual8.1 mslaya:en9.6 msgliclass14.7 msnli20.4 msdecider:0.8b155 msdecider:2b190 msTypeSafe Jevhosted API236–276 msOllaya: median of a five-question request through the HTTP API on an NVIDIA RTX 4090 (laya in fp16, the others in fp32). Jev: median request latency of the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), which includes the network. Setups differ, so read it as an order-of-magnitude comparison.Drop-in compatibleSpeaks TypeSafe's API.Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.Request# Point the TypeSafe SDK at Ollaya export TYPESAFE_BASE_URL=http://localhost:11435 export TYPESAFE_API_KEY=local # any value works export TYPESAFE_DEFAULT_MODEL=laya # …or call the compatible endpoint directly curl http://localhost:11435/v1/systemone -d '{ "model": "laya", "state": "Can I get an invoice for last month?", "questions": { "intent": { "type": "choice", "instructions": "What does the customer want?", "criteria": { "invoice": "Needs an invoice or receipt", "refund": "Wants money back", "other": "Anything else" } } } }'Response{ "model": "laya:en", "answers": { "intent": { "type": "choice", "choice": "invoice", "confidence": 0.9547, "probabilities": { "invoice": 0.9698, "refund": 0.0172, "other": 0.013 } } }, "usage": { "input_tokens": 43, "output_tokens": 0 } }TypeSafe compatibility guide Open modelsOpen weights, ready to pull.Start with Laya from Convai Innovations: an English model, a 100+ language model, a model fine-tuned for typed decisions, and a router that picks for you.Your data stays yoursPrivate by default.Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.PlatformsRuns where you work.A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers.

§2 Mixed · 49%

Every model runs on the CPU; an NVIDIA GPU on Linux, in WSL 2 or in Docker takes a request down to milliseconds.PlatformDesktop appCommand lineGPUmacOSApple siliconDesktop appMenu bar app.dmgCommand lineInstall scriptGPUCPU onlyWindows10 and 11, x64Desktop appDesktop app.exe or .msiCommand linePowerShell scriptGPUCPU onlyNVIDIA via WSL 2Linuxx86-64Desktop appDesktop appAppImage, .deb, .rpmCommand lineInstall scriptsystemd serviceGPUNVIDIA, CUDA 13LinuxARM64Desktop appNot availableCommand lineInstall scriptsystemd serviceGPUCPU onlyWSL 2Linux on WindowsDesktop appNot availableCommand lineInstall scriptSame as LinuxGPUNVIDIA, CUDA 13Dockeramd64 and arm64Desktop appNot availableCommand lineImage on GHCRGPUNVIDIA, CUDA 13:cuda image, amd64Install for your platform NVIDIA GPUs need driver R580 or newer; the installers fetch the CUDA libraries only when they find one.

§3 Mixed · 53%

On Apple, AMD and Intel GPUs, models run on the CPU.Get up and running in minutes.One binary, one command: ollaya run laya.DownloadmacOS, Windows, Linux and Docker · Apache-2.0 · GitHub