Open-Source Talorys Puts a Personal AI Agent on Cloudflare's Free Tier in One Command
TL;DR
Talorys is a new open-source, single-user AI assistant that deploys into your own Cloudflare account with one command, npx create-talorys@latest, and is built to stay inside the Workers Free plan. You get streaming chat, editable long-term memory, tasks, notes, projects, and reminders that fire from Durable Object alarms, so nothing has to stay online. Developer Roc Yu published the repo, the npm installer, and the Hacker News post within about an hour on October 10; at the time of writing the thread sits at 234 points and 118 comments, most of them arguing about whether "self-hosted" can mean "runs on Cloudflare." The interesting part for builders is not the chat box. It is the architecture: one private Worker, one SQLite-backed Durable Object holding everything, and a daily budget of 10,000 free Workers AI neurons that the entire design is shaped around.
What shipped
The GitHub repo was created at 09:46 UTC on October 10, the create-talorys installer landed on npm eight minutes later, and the Show HN went up an hour after that. The code is TypeScript under an MIT license, and the repo stands at roughly 360 stars and 21 forks less than a day in. Installer version 0.1.2 is current.
The pitch, from the README: "One person. One Cloudflare account. One command. One personal AI agent." The installer checks for Node 20.18 or newer, logs you into Cloudflare through Wrangler's normal OAuth flow, asks for an owner password, deploys a private Worker and a Pages project, runs health checks without spending any inference, and prints your pages.dev URL.
The architecture is the story
Here is the request path, straight from the project's architecture doc.
The browser only ever talks to a Cloudflare Pages site. Calls to /api/* hit a Pages Function, which forwards the original request over a service binding to a private Worker running a Hono router. That Worker is deployed with workers_dev: false and preview_urls: false, so it has no public URL at all; the only way in is through the binding. The Worker then calls getAgentByName with a fixed, server-side instance name, which means a client cannot address any agent other than yours.
That instance is a Durable Object built on Cloudflare's Agents SDK. It owns the entire database: conversations, messages, memories, tasks, notes, projects, automations, sessions, login attempts, and daily usage, all as SQLite tables inside the object. There is no D1, KV, R2, Vectorize, or Workflows anywhere in the stack, deliberately: none are needed and some cost money.
If you have not used Durable Objects: think of one as a tiny office with the filing cabinet bolted inside the room. Code and data live together, the lights go off when nobody is there, and the room is back in business the moment a request knocks. Scheduling rides on the same primitive. The Agents SDK's schedule API stores each pending task in a SQLite table and uses a Durable Object alarm to wake the agent when it is due, so a reminder set for Tuesday survives with no process running in between. Talorys keeps exactly one pending schedule per enabled automation, and on startup reconciles anything that was lost, running a job once if it was overdue by under 24 hours.
The free-tier math
The model is GLM-4.7-Flash on Workers AI, a 131,072-token-context model that Cloudflare lists at $0.06 per million input tokens and $0.40 per million output tokens. Cloudflare bills Workers AI in "neurons," its cross-model unit of GPU work, and the pricing page gives every account 10,000 of them a day at no charge, resetting at 00:00 UTC. Above that, on a paid plan, they cost $0.011 per 1,000.
Neurons are arcade tokens: the same coin buys a different amount of play depending on the machine. On GLM-4.7-Flash a million input tokens cost 5,500 neurons and a million output tokens cost 36,400, so the daily allowance works out like this.
Ten thousand neurons is worth about 11 cents at the paid rate, or roughly $3.30 a month of inference if you drained it every day. One HN tester reported burning about 500 neurons in a few minutes of chat and guessed the allowance covers 20 to 30 minutes of back-and-forth a day. Talorys is built around that ceiling: reasoning mode is switched off to save neurons, older history is summarized once it overflows the context budget, and the settings panel exposes caps for output tokens, context tokens, tool calls per request, AI requests per day, and scheduled AI runs per day. Plain reminders and the daily task digest never touch the model, so when the allowance runs out, chat pauses until midnight UTC and everything else keeps working.
The rest of the free plan is roomier. Durable Objects on the free tier allow 100,000 requests, 100,000 SQLite rows written, and 5 million rows read per day, with 5 GB of stored SQL data, and the Workers Free plan allows 100,000 requests a day at 10 ms of CPU each. For one person's assistant, inference is the only quota that will ever bind.
What it will not do
Read the tool list before you get excited. Tools only touch Talorys's own service layer: there is no shell, no code execution, no outbound HTTP, no web search, and no credentials. Scheduled AI runs get read-only tools plus a notify-owner tool. Destructive tools (delete a task or note, forget a memory, cancel an automation) require a confirmed: true flag and an explicit confirmation in your latest message, and the security doc stresses that this is enforced in code, not just in the system prompt. It is a notebook with a brain, not an agent that goes and does things on the internet.
Memory is the other soft spot. Recall is keyword search over SQLite FTS5 indexes, not embeddings, and the first substantive issue on the repo documents what that costs. A tester saved "I'm a vegetarian" as a memory, opened a new chat, asked for dinner ideas, and got Greek chicken bowls and salmon. The question shares no words with the memory, so the prefix search never finds it. A newer memory ("traded the Bug for a 2025 Tesla") also failed to override an older one that happened to contain the word "car." The free-tier decision to skip Vectorize is exactly what makes semantic recall awkward here, so expect that to be the first real fork in the roadmap.
The security posture is better than the average weekend project
The owner password is hashed on your machine with PBKDF2-HMAC-SHA256 at 100,000 iterations (the Workers maximum) and stored only as a Worker secret, handed to Wrangler over stdin. Sessions are 256-bit random tokens in an HttpOnly; Secure; SameSite=Strict cookie, and the database stores only an HMAC of each token, so a leaked database cannot be replayed. Five failed logins from one IP within 15 minutes lock that client, 25 failures lock login globally, sessions expire after 30 days or 14 days idle, and every non-GET request needs a custom header plus a matching Origin. The Pages site ships a strict CSP with script-src 'self' and frame-ancestors 'none', and the build checks that no secrets or bindings leak into browser code.
The "self-hosted" fight
Most of the 118 comments are a definitional brawl. The title says self-hosted; the runtime, storage, inference, and CDN are all Cloudflare. One commenter asked where to download Cloudflare. Another suggested the honest title was simply "an AI agent hosted on Cloudflare's free tier," and several settled on "self-managed" as the accurate term. The author's own framing is narrower than the title: no server, database, or account operated by the Talorys developers, and no telemetry, which is a real property even if it is not the one the word promises.
The useful rebuttal came from the people who read the code. The AI layer sits behind an AIProvider interface with generate, stream, and generateWithTools; the Workers AI implementation uses the AI SDK with Cloudflare's workers-ai-provider, and the doc says adding OpenRouter or Gemini means implementing that interface and selecting it in one method. A deterministic mock provider already exists for local development, so the whole stack runs on your laptop under wrangler dev with no Cloudflare login. Pointing it at a local model server is a small change. Calling the result self-hosted would, at last, be uncontroversial.
Caveats
- It is a day old. Version 0.1.2 and two open issues. Treat it as a reference design you can read, not infrastructure you depend on.
- Workers AI latency. One HN commenter called Workers AI "verrrrry slow for user interactive use cases," fine for background jobs. We have not benchmarked it; run a few turns before you commit.
- Quotas are Cloudflare's, not yours. The README says free-plan quotas "are set by Cloudflare and can change," and one commenter on a paid Workers plan reported being billed for neuron usage they believed was covered. If your account is on the $5 a month Workers Paid plan, overage is billed, so set a spend limit.
- GLM-4.7-Flash is the only production model. It is a fast multilingual model tuned for tool calling, not a frontier model. The free tier buys you a capable assistant, not a brilliant one.
- Keyword memory. See above. Do not trust it to remember your allergies.
Why builders should care
Strip the chat UI away and Talorys is a clean, copyable template for a single-tenant stateful service on Cloudflare that costs nothing at rest: a static Pages site, a Pages Function as the only public entry point, a service binding into a Worker nobody can reach directly, and one Durable Object that is simultaneously the database, the scheduler, and the agent loop. The installer is worth reading on its own: idempotent re-runs, collision-free resource names, secrets through stdin, health checks that skip inference, and a manifest audited for secret-shaped strings before it is written. If you have been meaning to learn Durable Objects, this is a better tutorial than most tutorials.
Key Takeaways
- Talorys deploys a single-user AI assistant with memory, tasks, notes, and alarm-driven reminders into your own Cloudflare account with
npx create-talorys@latest; MIT, TypeScript, roughly 360 stars on day one. - The stack is Pages plus a Pages Function, a service binding to a private Worker with no public URL, and one SQLite-backed Durable Object on the Agents SDK that holds all data and all schedules.
- The free tier's 10,000 Workers AI neurons a day buy up to 1.82M input tokens or 275K output tokens on GLM-4.7-Flash, worth about $0.11 a day at the paid rate; the design disables reasoning and summarizes context to stay under it.
- Tools are confined to the app's own data: no shell, no outbound HTTP, no web search. Scheduled runs are read-only, and destructive actions require code-enforced confirmation.
- Memory recall is FTS5 keyword search, and the first issue shows it missing a saved "vegetarian" preference; semantic recall would need the vector store the free-tier design avoids.
Sources: Talorys on GitHub, Talorys architecture doc, Talorys security doc, Talorys issue #2: proposal for semantic recall, create-talorys on npm, Hacker News discussion, Cloudflare Workers AI pricing, Workers AI model page: GLM-4.7-Flash, Cloudflare Durable Objects pricing, Cloudflare Workers pricing, Cloudflare Agents SDK: schedule tasks, Cloudflare Pages Functions bindings