No-Code Agentist
A daily digest of practical AI agent workflows for non-developers. We isolate actionable no-code guides from viral hype. Scored against human-defined standards.
Daily Summary
4 curated | 5 evaluatedThe no-code agent landscape is maturing rapidly, with builders moving beyond simple chatbots to autonomous systems that execute and . Discussions highlighted the gap between true agent capabilities—those with decision layers, planning modes, and self-correcting logic—and surface-level implementations. Meanwhile, pricing models and platform positioning debates continue, as builders distinguish between one-off assistance tools and genuinely autonomous, repeatable process automation.
Talos — Business in a Hermes Box For the Hermes Agent Accelerated Business Hackathon @NousResearch I built Talos — an autonomous business system that takes a company concept and produces a complete go-to-market operation. Strategy, product design, marketing, sales page, visuals. No human decisions beyond pressing Enter. What it does Talos is a 3-tier autonomous business stack running on a Raspberry Pi: Tier 1 — CEO Bot (Decision Layer): Takes a company brief and deploys research scouts to investigate the market. Evaluates product candidates, designs a complete sales funnel (FREE → $27 → $97 → $497/mo → $18K DFY), and defends every decision against a 5-lens critic panel. Outputs PRODUCT-SPEC.md — the source of truth for everything downstream. Tier 2 — Marketing Orchestrator (Execution Layer): Reads the CEO's spec and runs 13 marketing skills across 7 dependency waves: research → strategy → content → distribution → launch. Every output is verified by critics before the next step sees it. Produces sales page copy, email sequences, social posts, lead magnets, and a launch plan. Tier 3 — Self-Evolving Skills (Learning Layer): Every critic catch becomes a permanent rule. The system gets smarter with each cycle. The demo run One prompt. 82 minutes. 104 bots spawned autonomously. 5,248,710 tokens consumed, 50 web searches 94 files generated: 65 markdown docs, 10 HTML files, 9 visual assets, plus JSON configs, CSS, and an audit log 3 rounds of critic review, 2 fix iterations The bots designed a complete 5-tier product funnel, wrote all the copy, generated the visuals, and published a live sales page with a working Stripe checkout A real $97 purchase was made via Stripe on the page the bots built Three human interventions during the run — all gated approvals at critic checkpoints. Everything else was autonomous. The product the bots created Domination Engine™ — an AI Marketing Operating System. A decision layer that tells an established business which AI marketing move to make next, in what order, with tool-agnostic workflows to execute it. The CEO bot didn't just pick a product — it designed and priced the entire funnel, from free lead magnet to $18K done-for-you service. The PRODUCT-SPEC.md was detailed enough to hand directly to Claude Code to build the software. The sales page is live. The Stripe checkout works. Sponsor stack 🤖 Nous / Hermes — Hermes Agent orchestrates the entire 3-tier architecture. Dozens of skills, persistent memory, subagent spawning, profile isolation. CEO and Orchestrator run as separate Hermes profiles with distinct contexts. 🟩 NVIDIA Nemotron — Powers the reasoning across all bots. NemoClaw serves as the adversarial critic layer — the model that catches the other model's mistakes. 💳 Stripe — Live checkout on the bot-generated sales page. What makes this different Talos doesn't do a task. Talos takes a business concept and produces a complete go-to-market operation — researching the market, making product decisions, designing a funnel, writing copy, generating visuals, publishing a sales page, and collecting revenue. The CEO bot doesn't just execute instructions. It makes strategic decisions and defends them against critics. The system runs on Windows 11 utilizing Hermes. Orchestration is local. Inference runs on NVIDIA's cloud. Stripe used for checkout. Bots that catch their own bullshit. Links 🌐 Live sales page: https://t.co/RGN31igwd4 💻 Full output (94 files): https://t.co/ZPDcilcvrq
99% OF PEOPLE ARE BUILDING AI AGENTS WITH CLAUDE THE WRONG WAY They chat with it and call it an agent.A real agent has structure, knows when it doesn’t know something, and asks before taking action. In this no-code tutorial, I show you how to build actual working AI Agents using Claude Code — only with Claude Desktop and a folder on your computer. Inside the video: The 3 levels of Claude (and why most people never go beyond level 1) Setting up your workspace the right way Why CLAUDE.md changes everything How to use Planning Mode so Claude stops guessing Building two real agents live The 5 deadly mistakes that destroy agent workflows Stop prompting. Start building.Chapters below. CC: Ai Master https://t.co/6p3ThgbsPH
The AI market still prices models per token, but users experience them per accepted result. Once smaller models compensate for lower intelligence by thinking longer, generating more, retrying more, or needing more human correction, they can become pseudo-cheap rather than actually cheap. That is the hidden insight. OpenAI officially describes GPT-5.6 Sol as the flagship tier, Terra as the balanced everyday tier, and Luna as the fast/affordable tier; pricing is $5/$30 per million input/output tokens for Sol, $2.50/$15 for Terra, and $1/$6 for Luna. OpenAI also says Terra is competitive with GPT-5.5 at 2x lower cost, and that GPT-5.6 is currently in limited preview through API/Codex for selected organizations, not general ChatGPT availability yet. Anthropic says Sonnet 5 is the default model for Claude Free and Pro, with introductory API pricing of $2/$10 per million input/output tokens until August 31, 2026, then $3/$15; Opus 4.8 is $5/$25. So the facts support a sharper version of your concern, but a few claims need tightening. The best thesis Current draft thesis: Small models seem strictly worse than larger models at lower reasoning effort. Stronger thesis: The new frontier is not “small model versus big model.” It is “cheap token versus cheap completed task.” Small models can look cheap on the pricing page while being expensive at the workflow level because they spend more tokens, require more retries, and lose the compressed judgment you get from frontier models. Even stronger: A lower price per token is not a lower price per unit of cognition. Best version: The old model market was about price per million tokens. The new model market is about dollars per accepted result, seconds per accepted result, and interventions per accepted result. On that axis, some “cheap” models may be quietly dominated by larger models at lower reasoning effort. That is the money line. Claim hygiene: what to fix Draft claimProblemStronger version“Terra and Luna seem strictly worse than Sol on lower reasoning effort.”“Strictly worse” is a mathematical claim. You need task distribution, latency, retry, and success data.“On many complex workflows, Terra/Luna may be dominated by Sol-low once you price in token bloat, retries, and quality loss.”“Sonnet 5 is worse than Opus 4.8 at the same price points.”Anthropic claims Sonnet 5 offers better cost-performance over a wider range than Sonnet 4.6 and can match Opus 4.8 on some tasks. You need to challenge the methodology, not ignore it. “The question is whether Sonnet 5’s cost-performance curve holds under real workflows that measure accepted outputs rather than benchmark pass rates.”“Sonnet 5 on the free tier may cost Anthropic more than just giving users Opus.”This is only true if Sonnet uses enough extra tokens/retries to overcome Opus’s higher price. At intro pricing, Opus is 2.5x Sonnet’s token price; after September 1, it is about 1.67x. “Sonnet only costs more than Opus if its workload-level token multiplier plus retry tax exceeds the Opus price premium.”“Small models are faster.”Sometimes true per token, but not necessarily per task.“Per-token speed is a component; wall-clock-to-correct-answer is the metric.”“Per-token measurements are useless in 2026.”Too absolute.“Per-token measurements are no longer sufficient. The useful metric is quality-adjusted wall-clock goodput.”“Why would they release these models?”Strong question, but answer is not only intelligence/price.“Providers may release dominated-looking tiers for capacity shaping, safety gating, free-tier economics, fleet utilization, quota design, latency distribution, and router architecture.” The killer concept: token price is a decoy The line you want: In 2026, the unit of AI work is not the token. It is the accepted artifact. A token is a billing unit, not a productivity unit. The real unit might be: one merged pull request, one verified research answer, one resolved customer ticket, one working spreadsheet, one correct simulation, one successful agent run, one decision that survives expert review. So the proper metric is not: cost per 1M tokens It is: cost per accepted result = tokens + retries + tool calls + latency + human correction + failure risk This is the central missing element. Break-even math that makes the argument rigorous Sonnet 5 versus Opus 4.8 At launch pricing, Sonnet 5 is $2 input / $10 output per million tokens, while Opus 4.8 is $5 input / $25 output. That means Opus is 2.5x more expensive per token during the Sonnet 5 intro window. After September 1, Sonnet 5 becomes $3/$15, so Opus is about 1.67x more expensive per token. So this claim: “Sonnet 5 free tier might cost Anthropic more than Opus” only holds if Sonnet’s real workload cost multiplier exceeds those thresholds. More precise: During the intro period, Sonnet 5 would need to use more than 2.5x as many effective billable tokens as Opus 4.8 for Opus to be cheaper on raw API pricing. After September 1, the break-even falls to about 1.67x. But “effective tokens” should include: output tokens, hidden/thinking tokens where billed, retries, failed attempts, context drag, compaction, tool-call overhead, user correction loops, safety interruptions, rerouting to a bigger model. The better version: Sonnet does not need to be cheaper per response. It needs to be cheaper per completed job. That is a much harder bar. Sol versus Terra/Luna OpenAI’s GPT-5.6 pricing gives a clean break-even: Sol: $5 input / $30 output Terra: $2.50 input / $15 output Luna: $1 input / $6 output So Terra is half the token price of Sol. Luna is one-fifth the token price of Sol. That means: Terra only beats Sol economically if it can do the task using less than 2x Sol’s effective token/retry budget. Luna only beats Sol if it can stay under 5x. The obscure but crucial point: a lower-effort big model can have higher cognitive compression. It may produce shorter plans, fewer dead ends, fewer clarifications, and fewer repair attempts. That can erase the smaller model’s price advantage. Fable 5 versus Sonnet 5 Anthropic lists Claude Fable 5 at $10 input / $50 output per million tokens, while Sonnet 5 is $2/$10 during intro pricing and $3/$15 afterward. So Fable is: 5x Sonnet 5’s intro price 3.33x Sonnet 5’s standard price Therefore: Fable must use less than 20% as many effective tokens as intro-priced Sonnet, or less than 30% as many effective tokens as standard-priced Sonnet, to win on raw token cost. But it can still win if it dramatically reduces failed runs, human intervention, or time-to-correct-artifact. The sophisticated phrasing: The larger model does not have to be cheaper per token. It only has to be sufficiently more compressed per unit of solved work. The phrase you need: semantic compression “Big model smell” is real, but make it more legible. Use: semantic compression Meaning: the big model compresses more task understanding into fewer visible moves. A larger model may: infer the actual goal from messy instructions, avoid pointless intermediate steps, choose better abstractions, recognize when a benchmark-style answer is insufficient, maintain global coherence over long tasks, detect traps before spending tokens on the wrong path, know when not to over-explain, ask fewer unnecessary questions, recover from tool failures better, produce artifacts that need less human cleanup. That is the hidden advantage benchmarks often miss. Better than “big model smell”: Benchmarks measure answer correctness. Real workflows also depend on taste, compression, recovery, calibration, and not wasting the operator’s attention. Or: The frontier model advantage is not just higher IQ. It is lower entropy per decision. The missing framework: price-per-token versus cognition-per-token Add this distinction: Cheap tokens A model is cheap because each token costs less. Efficient cognition A model is cheap because it needs fewer tokens, turns, retries, and corrections to reach the accepted answer. Real productivity A model is cheap because the human gets the result faster with less supervision. Most AI pricing discourse stops at the first. Your post should be about the gap between all three. The line: There are cheap tokens, cheap thoughts, and cheap outcomes. These are not the same thing. Why providers still release “dominated-looking” models This is the most important missing section. The draft asks “why would they release them?” but does not give enough possible answers. 1. Fleet utilization A model can look dominated on user-facing API prices while being rational internally because it uses different hardware, different serving paths, different batching behavior, or different capacity pools. A provider may have excess inference capacity that is poorly suited to the largest model but perfect for a smaller tier. The public price is not the same as internal cost. The line: The public Pareto frontier is not the provider’s fleet Pareto frontier. 2. Queueing economics Even if a small model has worse average cost per solved hard task, it may improve p95 latency for simple traffic by keeping trivial requests away from the expensive frontier fleet. The hidden metric is not average speed. It is: congestion avoided per request routed away from the flagship model. 3. Safety and capability gating Free-tier users may not get the strongest model by default because of misuse risk, not just cost. Anthropic says Sonnet 5 has a much lower ability to perform cybersecurity tasks than current Opus models, while still being safer in agentic contexts than Sonnet 4.6. That alone is a plausible reason to prefer Sonnet as a default for broad/free access. The line: A model can be “worse” in capability and better as a default product because it has a safer capability envelope. 4. Subscription quota design In consumer products, the binding constraint is often not API cost. It is abuse, rate limits, perceived fairness, and retention. A free user does not behave like an API customer. They may send short prompts, low-stakes questions, and many abandoned conversations. A cheaper-enough default model may be optimal even if it loses on expert coding benchmarks. 5. Router architecture Small models are not always meant to be the final answerer. They can be used as: classifiers, routers, summarizers, context compressors, draft generators, tool-call planners, cheap verifiers, first-pass executors, fallback models, abuse filters. A model that is dominated as a standalone assistant can still be useful inside a compound system. The line: The question is not “Would I manually choose Sonnet over Opus?” The question is “Where does Sonnet improve the router?” 6. Market segmentation Providers need visible tiers. Not because every tier is globally optimal, but because customers have different budget psychology, procurement rules, latency requirements, risk tolerance, and feature access constraints. A “balanced” model can be commercially useful even if power users can find a better frontier-model-low-effort arbitrage. 7. Availability A model that is 10% worse but always available can beat a model that is better but capacity-constrained. The real enterprise metric is not benchmark score; it is service-level reliability under load. 8. Model behavior preference Some users prefer smaller models because they are less intense, less agentic, less verbose, less likely to over-engineer, and easier to steer for routine tasks. That is not intelligence. It is temperament. 9. Data and training flywheel Broad defaults generate interaction data, product telemetry, and failure cases. Even when data is not used directly for training in some paid settings, aggregate product signals can still inform evaluation, routing, prompt design, and future model development. 10. Strategic anchoring A lower tier makes the flagship feel premium. It creates a ladder: free/default → pro/balanced → flagship. This is product architecture, not just inference math. The obscure thing almost everyone misses: tokenizers corrupt comparisons Raw token counts across model families are not apples-to-apples. Anthropic explicitly notes that Claude Opus 4.7+ models, Fable 5, Mythos 5, and Sonnet 5 use a newer tokenizer that produces approximately 30% more tokens for the same text; Anthropic’s Sonnet 5 post says the same input can map to roughly 1.0–1.35x as many tokens depending on content type. So your argument needs to split token bloat into three categories: Tokenizer bloat Same text becomes more tokens because of tokenizer changes. Reasoning bloat The model needs more intermediate thinking/output to solve the same problem. Workflow bloat The model causes more retries, clarifications, repairs, and human edits. Only #2 and #3 are true cognitive inefficiency. #1 is partly accounting. The line: Not all token inflation is stupidity. Some of it is tokenizer accounting. The real question is effective tokens per accepted result. The “per-token speed is useless” point, refined Your draft says: “Per-token measurements are utterly useless in 2026.” Make it sharper: Per-token speed is a component metric, not a user metric. Users experience wall-clock-to-correct-artifact. A model can be “fast” at 300 tokens/sec and still feel slow if it emits 10,000 tokens, wanders, retries, or forces the human to read a giant answer. Google says Gemini 3.5 Flash supports a 1M-token context, 65k max output, thinking, and computer-use features, and Google positions it as highly speed-oriented; its API pricing is $1.50 input / $9 output per million tokens for standard paid use. Google also says Gemini 3.5 Flash is 4x faster than other frontier models by output tokens per second. That does not settle user experience. Better metric: TCA: time to correct artifact. Formula: TCA = time to first token + generation time + tool-call time + safety/classifier delay + retry time + human inspection time + human correction time Even better: Goodput = accepted artifacts per dollar per hour. That is the 2026 metric. The missing benchmark design To prove your claim, the benchmark must not be “which model answered correctly once?” It should be: Given the same dollar budget and the same wall-clock budget, which model produces the most accepted outputs? Measure these pass@1 pass@$1 pass@minute accepted-artifact rate total tokens including thinking where billed retry count tool-call count human intervention count human review time p50/p95 wall-clock latency refusal/safety-interruption rate compile/test success for code edit distance from accepted final answer number of turns until done output verbosity context compaction frequency cache hit rate hallucinated action rate “looked done but wasn’t” rate The last one is important. Small models often fail in a particularly expensive way: they produce something that looks plausible enough to require human inspection but not correct enough to ship. Call this: plausible slop tax. The “big model smell” rubric Turn the vibe into a checklist. A model has “big model smell” when it does these things: Reads the room It infers the real objective behind the literal prompt. Compresses the problem It finds the right abstraction instead of grinding. Avoids premature closure It does not declare victory after a shallow solution. Maintains global state It remembers the purpose of earlier decisions in long workflows. Handles contradictory constraints It negotiates tradeoffs instead of ignoring one. Recovers from tool failures It changes strategy rather than looping. Knows when not to think more It does not burn effort on easy subtasks. Produces fewer fake certainties It distinguishes “probably” from “verified.” Has taste It chooses solutions that are elegant, maintainable, or contextually appropriate. Reduces operator burden It makes the human feel like a reviewer, not a babysitter. That is what benchmarks undercount. The line: The big model advantage is often not the answer. It is the absence of babysitting. The better name for the phenomenon Use one of these: 1. Token inflation tax Cheap models become expensive when they inflate the number of tokens required to reach an acceptable result. 2. Cognitive compression premium Frontier models are expensive per token but cheap per unit of compressed reasoning. 3. The false economy zone The region where a model is cheap enough to attract use but not strong enough to reduce total workflow cost. 4. The pseudo-cheap model problem A model that wins the pricing table and loses the task. 5. Reasoning arbitrage Using a larger model at lower effort to beat a smaller model at higher effort. My pick: reasoning arbitrage It is memorable and precise. The hidden Pareto frontier Your original “Pareto frontier” point is strong, but broaden it. A model is not on one frontier. It is on several: cost versus accuracy, cost versus latency, cost versus safety risk, cost versus availability, cost versus context length, cost versus human supervision, cost versus retry rate, cost versus refusal rate, cost versus enterprise compliance, cost versus fleet capacity. A model that is dominated on one visible frontier may be useful on another hidden frontier. The line: What looks dominated on the public benchmark frontier may be useful on the provider’s hidden fleet frontier. That is the fairest steelman. Stronger version of your argument Here is a much sharper rewrite: My concern with the new “middle” model tiers is that they may be pseudo-cheap.The pricing page says they are cheaper per token. But users do not buy tokens. Users buy completed work.That is why GPT-5.6 Terra/Luna and Claude Sonnet 5 are interesting. The question is not whether they are good models. They clearly are. The question is whether they occupy a real cost-performance frontier once you compare them against larger models running at lower reasoning effort.If Sol-low solves the task in fewer turns, fewer retries, and fewer total tokens than Terra-high, Terra is not cheaper. It is just cheaper-looking.Same with Sonnet 5 versus Opus 4.8. Sonnet’s list price is lower, but if it needs more reasoning, more output, more retries, and more human correction, the effective price gap can collapse. And benchmarks often miss the “big model smell” advantage: taste, compression, recovery, calibration, and not making the operator babysit the run.The old metric was dollars per million tokens.The new metric is dollars per accepted artifact.This is why per-token speed has become such a weak benchmark. A model can generate tokens quickly and still be slow in practice if it thinks too verbosely, loops, or creates more text than the user can inspect. Wall-clock-to-correct-result matters more than tokens/sec.The steelman for Anthropic and OpenAI is that these models may be useful for hidden reasons: free-tier safety, fleet utilization, rate-limit design, capacity shaping, routing, simple-agent execution, and lower-risk default behavior. A model can be dominated as a manual choice and still be valuable inside a product router.But as a user, I care about the visible economics:Which model gives me the correct artifact with the least total cost, least wall-clock time, and least supervision?On that metric, the uncomfortable possibility is that many “cheap” models are not cheap. They are token-discounted but cognition-expensive.The real benchmark should not be price per https://t.co/6lFgR6rRxD should be:accepted results per dollar per hour. Punchier version Small models might be entering their false economy era.The market still prices AI by the token, but serious users experience AI by the completed artifact.That distinction matters. A model can be cheaper per token and more expensive per task if it thinks longer, talks more, retries more, fails more, or needs more human cleanup.This is my concern with tiers like GPT-5.6 Terra/Luna and Claude Sonnet 5. They may look cheaper on the invoice but lose to bigger models at lower reasoning effort once you measure total workflow cost.The hidden variable is cognitive https://t.co/RDjCJKiotV models often do not just know more. They waste less motion. They infer the real objective, avoid bad branches, recover from tool failures, and produce artifacts that need less babysitting.That “big model smell” rarely shows up cleanly in benchmarks.The steelman is that smaller models still make sense for providers: capacity shaping, free-tier defaults, safety gating, routing, simple automation, and fleet utilization. A model can be commercially rational even if it is not what a power user would manually choose.But for users, the metric is simple:Do not ask what the token costs. Ask what the accepted result https://t.co/4JTgZoNzu5 2026, tokens/sec is a vanity metric.The real frontier is:correct artifacts per dollar per hour. Even more aggressive version A lot of “cheap AI” is not cheap. It is just metered wrong.Per-token pricing made sense when models mostly answered questions. It breaks down when models act like https://t.co/QHNaKhB8b5 agentic work, the cost is not the token. The cost is the loop: planning, tool calls, retries, failures, compaction, verification, and human supervision.That is why some smaller models may be structurally fake-cheap. They cost less per token, but they spend more tokens to think worse.A big model at low effort can sometimes beat a small model at high effort because intelligence compresses search. It takes fewer wrong branches. It needs fewer clarifications. It produces less plausible slop. It knows when the job is actually done.This is the part benchmarks miss. They measure correctness, but not taste. They measure pass rates, but not babysitting. They measure token speed, but not time-to-trust.The real unit is not the token.The real unit is the accepted artifact.Until model rankings show cost per accepted artifact, the pricing tables are theater. The best one-liners Use these anywhere: Cheap tokens are not cheap cognition. The unit of AI work is shifting from token to accepted artifact. A model can win the pricing table and lose the workflow. The danger zone is token-discounted but cognition-expensive. Frontier models can be expensive per token and cheap per decision. Per-token speed is not UX. Wall-clock-to-correct-artifact is UX. The smaller model has to beat the bigger model’s semantic compression. A low-cost model that needs retries is just deferred cost. The frontier is accepted results per dollar per hour. The big model advantage is not always intelligence. Sometimes it is less babysitting. Token/sec is a component metric. Goodput is the product metric. If the model makes the human read more, it is not cheap. Small models are often cheaper at generation and more expensive at supervision. The hidden benchmark is time-to-trust. Some models are not cheap; they are just low-denomination. Obscure but useful angle: “human attention” is the unpriced input The biggest missing variable is not tokens. It is human attention. A small model that writes 4,000 plausible tokens may be more expensive than a big model that writes 900 high-signal tokens, even if the API bill is lower. Why? Because human review is the scarce resource. Add this: The most expensive token is the one the human has to inspect. Or: A verbose cheap model can transfer cost from GPU time to human time. This is especially true for: code review, legal drafting, medical research, finance analysis, scientific reasoning, simulation debugging, strategy work, long-horizon agents. The model does not just generate output. It generates review burden. Obscure angle: “plausible wrong” is more expensive than “obviously dumb” Smaller models often fail expensively because they are good enough to sound right. That creates: verification debt A weak model that obviously fails is cheap to discard. A mid model that sounds plausible requires inspection. The line: The most dangerous model tier is not the dumb one. It is the one that is articulate enough to create verification debt but not strong enough to earn trust. That belongs in the post. Obscure angle: “effort labels are not comparable” “High,” “medium,” “xhigh,” “thinking,” and “max” are product labels, not universal units. OpenAI says GPT-5.6 introduces a new max reasoning effort for Sol and an ultra mode using subagents; Anthropic’s Sonnet/Opus ecosystem uses effort/adaptive thinking in its own way. So instead of comparing “high effort” across vendors, compare: same dollar budget, same time budget, same acceptance criteria. Better line: Effort labels are marketing coordinates. Dollar-normalized performance is the real coordinate system.
What it is: A no-code platform for building AI agents — automated workflows where AI executes a defined set of steps repeatedly, triggered by conditions you set. You build the agent once. It runs as many times as needed without your involvement. What it is NOT: It's not a chatbot. It's not a writing assistant. It's not a tool you "use" in the traditional sense. If you're looking for something to help you draft emails or write captions — Rytr or AiAssistWorks fit that job better and cost less. Relevance AI is for automating repeatable multi-step processes. Pricing: A free plan is available with limited credits. For paid plans, the best way is to talk to their sales directly and get exactly what you need.