No-Code Agentist
A daily digest of practical AI agent workflows for non-developers. We isolate actionable no-code guides from viral hype. Scored against human-defined standards.
Daily Summary
12 curated | 20 evaluatedSeptember 4 spotlighted the tension between agent ambition and practical delivery: with promises to identify and fix project issues autonomously, while practitioners documented the gap between , and Microsoft shipped to help low-code makers build multi-step business agents. Beneath the platform announcements, builders confronted hard limits: context budgets for agent instructions remain , forcing painful tradeoffs in what agents can see and use.
GROK BOT HOW TO USE WHAT NEW TODAY - Folder now 840 posts (26 new since Sep 2) - Official (bot): Grok Bot is now available on Android. https://t.co/ivi4xRy86g - Official Elon: Your Bot will identify issues to resolve, notify you when they’re fixed and continue on with your project https://t.co/l3wz7YZgCI - Testers (mark_k): Grok Bot for Android is here! You can install Grok Bot right now from the Play Store. Everything works as expected, bringing the full Grok Bot experience to A… https://t.co/AX6hUUQ54y - Testers (RoundtableSpace): SpaceX AI is building a Grok Bot Marketplace where you browse specialized bots by category, see exactly how each works and add them to your team with routines… https://t.co/wpzcGOAaXr - Testers (farzyness): This is actually insane. Yesterday I asked the X team to fix the API so that I can upload videos longer than 10 minutes. My Grok bot JUST pinged me that this… https://t.co/uZeoWUkFkd - Testers (kloss_xyz): Stop what you're doing. I read all 5 official Grok Bot guides. 11 pro tips barely anyone’s talking about: 1. Instead of telling it your priorities, have it rea… https://t.co/zHFiDEDzRF - Testers (s4yonnara): Claire Vo just walked through how she runs her entire business and family on a team of named Grok Bots, one hour, beginner to pro: "I hire my bots like I hire… https://t.co/7YVobhOxoe - Testers (adiix_official): SpaceXAI team just dropped a 3-page operator’s manual for turning Grok Bot into a full multi-agent system that runs entire workflows 24/7 The shift: instead of… https://t.co/anck7exuJO - Testers (unicodef1wn): Most people prompt Grok Bot like they prompt ChatGPT. It's built to be a teammate, not a chat window. Most people open Grok Bot, type a vague task, and wonder… https://t.co/g7FqLfHXxy - Testers (joehansen): The Grok Bots never sleeps! Elon already described the human version: wake up, work, sleep, then do it again seven days a week. A Grok Bot is that shift withou… https://t.co/VD7pZCGXpI - Testers (tibor_tee): Grok Bot for Product: Best Practices is Wednesday, Sep 2, 5-6pm UTC. Kevin Niparko and Roshan Sadanani show how to build agent teams to research, prioritize, a… https://t.co/zo4Ohf2zN6 - Testers (0xCarnagee): SpaceXAI engineer, Dawson Lind: "GrokBot and Cursor are the most powerful agentic tools we have ever shipped, but almost nobody runs them the way they were bui… https://t.co/qVaj6Py9Nm - Testers (0xRafy): SpaceXAI engineer, Lauren Tan: "GrokBot is the most powerful agentic tool we have ever built, but only 1% of users use it correctly right now I'm running a tea… https://t.co/o9cRyk8yAe - Testers (0xrux): Grok Bot just came out and I’m honestly blown away in just 5 minutes of testing it. her words: “it feels like Hermes or OpenClaw, but faster and easier. let me… https://t.co/QGP01LgPG8 WHAT IT IS - A Bot is a named, persistent teammate, not a one-off chat - You message it like a coworker and it finishes work in real apps - All of your Bots share one cloud computer (browser, files, terminal) - Each Bot has its own screen on that shared computer - Work keeps going when your laptop or the app is closed - It only comes back when it needs approval or a human step - Do not confuse Grok Bot with Digital Optimus (real-time video human emulation) - Testers (joehansen): do not confuse Grok Bot with an X Chat Agent (an X account you add to a group chat). Grok Bot is the machine that does the work - Official Elon: Your Bot will identify issues to resolve, notify you when they are fixed, and continue on with your project https://t.co/l3wz7YZgCI ACCESS AND INSTALL - Official bot (Aug 26): All SuperGrok and Cursor Pro subscribers now have access - Testers (mattyp): Grok Bot is free to try; every Grok & Cursor plan includes Grok Bot https://t.co/5RUqLmOoO9 - Testers quoting SpaceXAI (cb_doge): SuperGrok, SuperGrok Plus, SuperGrok Heavy, Cursor Pro, Cursor Pro+, Cursor Ultra, Cursor Teams Standard and Premium - Cursor CEO Michael Truell: everyone with a standard Grok or Cursor subscription - Testers: included at no extra cost with those plans; circulating plan prices $20/month Cursor Pro and $25/month SuperGrok (nextbigfuture) - Official bot + Elon: weekly / free usage limits reset for all Grok Bot users - Sign in with your Cursor account (that account owns plan and usage) - Official desktop download: https://t.co/tHvaZ7hExT - Also listed at https://t.co/bfJPYFHea1 - Official desktop is macOS and Windows (Apple silicon or Intel; Windows x64 or Arm64) - Testers (mattyp): Linux is now supported (official docs may lag) https://t.co/5RUqLmOoO9 - iPhone: iOS 18+, App Store app "Grok Bot" - Official (bot, Sep 2): Grok Bot is now available on Android https://t.co/ivi4xRy86g - Testers (mark_k): install from Play Store; full Grok Bot experience on Android https://t.co/AX6hUUQ54y - Play Store: https://t.co/o7XKDJuomX - Official docs may still lag on Android; iPad still not supported for actual use per older Official docs - SuperGrok Heavy users can choose Get access with SuperGrok Heavy, then Link Grok Account - Linking SuperGrok Heavy can unlock free Cursor Ultra (one Grok account to one Cursor account) - Legacy Privacy Mode blocks Grok Bot; switch to a supported Cursor data setting first - Grok Bot checks for updates automatically; Check for Updates is in Settings → Beta - Bookmark testers: iPad may be coming soon; official docs still say iPad is not supported yet SIGN IN [] Desktop: Get started, then Sign In with Cursor in the browser [] iOS: Login with Cursor, finish in the browser, return to the app [] Official: Cmd-D turns on voice dictation [] Official: it is easier to use the remote computer from your phone [] Use the same Cursor account that should own usage [] First-run tour asks which tools you use; that only shapes teammate suggestions [] Computer setup runs in the background, then Meet a future teammate opens CREATE A BOT [] New in the sidebar, or Cmd/Ctrl+N [] In New chat, choose Create new agent [] Edit Profile: name, title, description, avatar [] Give a short name, one primary job, and how it should work [] Description = standing rules (never send without approval) [] Chat messages = this-task instructions (draft follow-ups for these 12 accounts) [] Focused Bots beat one catch-all General Helper [] Testers (XFreeze): do not design a Bot from scratch; ask a Bot to create a capable one for you (it already has your context) [] Bookmark testers: start with about six roles; give each Bot recurring work [] Pin important Bots; hide unused ones (hiding does not pause routines) [] Bookmark testers: resize the sidebar, give titles, and group Bots into sections [] Duplicate a Bot to reuse a role for a new scope (copy does not include memory or history) [] Delete only if you are sure; files and logins on the shared computer stay [] Account limit: 50 Bots and group chats combined [] Name language can steer reply language (bookmark testers reported this) [] Bookmark testers: no model picker yet; testers still list voice mode, BYOK, multiplayer, hang-error messages, and a real iPad app as missing [] Testers (XEthanai / XFreeze): shareable Bot templates - build once, fine-tune, share the full setup for others to use or customize https://t.co/GCFpGeLso8 [] Testers (XFreeze): Marketplace coming - browse specialized Grok Bots by category and add them to your team https://t.co/h5VJjuPp2S [] Testers (RoundtableSpace): Marketplace framing — hand-picked bots by category with routines/workflows built in (app store for AI teammates) https://t.co/wpzcGOAaXr [] Testers (kloss_xyz): 8 official-ish bot templates (calls, X, writing, images, coding, networking, tool testing, parking tickets) https://t.co/R0xoBVRWHF HOW TO WRITE A TASK [] Outcome: what should be finished [] Sources: which apps, sites, files, or chats matter [] Constraints: what it must not do, or must ask first [] Deliverable: what it should return [] Review point: when it should stop for you [] Safe first task: attach a file and ask for a cited summary; do not change the file [] Next task: one real tool, read-only, with a sign-in handoff if needed [] After a good result, name the lasting format preference in chat [] Then save it as a skill or routine [] Official bot examples: email cleanup, Starlink-likely flights, build a site then buy the domain and deploy, meeting notes, sales prospecting, refunds, podcast audio digest [] Testers: paste a viral video link and ask the Bot to recreate it [] Testers: start from an org chart or make one Bot the boss of a repo, not a long to-do list [] Testers: one clear job, explicit allowed sources, no-guessing rule, approval gate before any real action [] Testers: 5 calibrations - confidence threshold, checkpoint long tasks, show negative examples, gate destructive actions separately, test edge cases first [] Testers: one Bot, one job - two jobs claimed worse at both; put recurring work on a schedule [] Testers quoting claimed SpaceXAI employee: version prompts and tool configs like code; scratchpad separate from the user-facing answer; retry with a different strategy; measure cost per successful outcome; review failure logs on a schedule [] Testers (mvanhorn): Compound Engineering - force a plan first; three-layer instructions (base, role, current focus); screenshot instead of retyping what's on screen; ping only on a real hit [] Testers (AlexFinn): anything you are about to do on your computer, ask Grok Bot first [] Testers (AlexFinn): example week - Reddit microSaaS, 3D digital office to watch Bots, Omarchy setup, product returns, daily AI news post, Cursor cloud-agent bug fixes, X model-release watch, brand-deal negotiation, contract redlines [] Testers (Michael_Fenech_): give it business responsibilities (checking, chasing, updating, coordinating), not just questions AGENT COMPUTER AND LOGINS [] Open Agent Computer from the conversation to watch the desktop [] Take over for password, passkey, 2FA, CAPTCHA, payment, or human-only pages [] Complete only the blocked step, then return control [] Never paste passwords or one-time codes into chat [] Use a secure secret request when the app offers one [] Browser sessions persist and are shared across all your Bots [] Prefer a connector/plugin when one exists; use the browser when it does not [] Testers (SamSokolin): for repeated web-app clicks, reverse-engineer browser-use into scripts (capture network once, call APIs next) https://t.co/bG9VoSeYuu [] Keep durable files in /workspace with clear project folders [] Settings → Beta: Update Agent Computer (preserve state), Recover, or Reset (can lose recent work) [] The cloud computer is not your Mac or Windows machine [] Local-computer execution is separate: Settings → General → Agent → Execution on Local Computer [] Default local policy is Ask every time; use Never allowed unless you need local files [] Bookmark testers: give a Bot full local Mac/iMessage access only on a spare machine [] Testers: local Mac access can write local Python scripts on your Mac [] Testers (Av1dlive / mvanhorn): Claude Code, Codex, and Cursor claimed runnable on the Bot's own computer [] Testers (XFreeze): take over that same cloud computer from your phone without being at the laptop [] Testers (Damir_Akaza): group-chat Bots can reach a local Mac Mini via Tailscale for project files and local scripts [] Testers (farzyness): a Bot can exclusively use Grok Build in CLI with the latest Grok models at highest thinking [] Testers (AlexFinn): a 3D digital office on a second monitor to watch Bots work PLUGINS AND CONNECTORS [] Settings → Plugins, then Add, then authenticate in the browser [] Grok Bot team (poteto): tinkabot helps build MCP/skills plugins and submit them for approval so all users can use them https://t.co/HFhRc1OGfa [] Type @ to attach a connector; type / to reference a saved skill [] Installed connectors are account-wide, not isolated to one Bot [] Official: you can connect multiple accounts to the same plugin [] Official bot (Aug 26): SuperGrok and Cursor Pro now included; weekly limits reset. Testers quoting SpaceXAI: also Plus/Heavy/Pro+/Ultra/Teams [] Testers: give a Bot its own email address so it can send and receive on its own [] Testers (mvanhorn): give a Bot its own inbox, not your Gmail (one mistake should not flag your domain); Twilio number so a Bot can make calls [] Testers: app v0.23.0 adds creating new channels [] Official bot: improved X support - connect your X profile; a developer account is auto-created with included credits https://t.co/NBucim3AGw [] Testers (blankspeaker): Plugins, add X, authenticate; $100 in X API credits good for a year (new and existing devs) [] Testers (GrokInsider): if credits did not land, ask Grok Bot to connect X so the Authorize card opens the browser; tester's manual plugin add failed [] Follow the Connect card when a Bot asks for a plugin [] Bookmark testers: X connector may still need a paid X API key in some setups (official path now auto-creates a developer account with credits) [] Bookmark testers: image generation works; testers now also claim video editing (music sync). Video generation still unconfirmed. Testers (XFreeze): dedicated video-editor Bot for footage → clips/cuts/sound/titles [] Bookmark testers: Grok Bot can now read your X bookmarks [] Testers (testerlabor): you can link your X Premium+ subscription to Grok Bot [] Testers: Higgsfield x Grok Bot; 100 free credits for new users [] Testers: send a YouTube link with start/end timestamps and get an HD clip in chat (Google login claimed) [] Testers: give a Bot Grok Build + Cursor CLI so weekly limits can be spread across Cursor and Grok [] Grok Bot team (poteto) + testers (XFreeze): Microsoft connectors live - Outlook, Outlook Calendar, OneDrive (search/send email, drafts, schedules, meetings, browse/upload files) https://t.co/c7HBr7lBXs https://t.co/pkvAZjDCSI [] Testers (XFreeze): Bot can natively generate images with Grok Imagine Image 2.0 inside the workflow (no separate Imagine tab) https://t.co/JKDh3dQo2q SKILLS AND ROUTINES [] Skill = reusable how-to (steps, rules, output, approval boundary) [] Routine = when to run that work (schedule or event) [] Do the task once, make it reliable, save a skill, then automate [] Teach a task: open computer view, choose Teach a task, demo up to 10 minutes, review the draft skill [] Teaching records the screen, not microphone audio; do not expose secrets while teaching [] If Teach a task is missing, ask the Bot to write a skill from the completed work [] Create a routine on the Bot that should own the job [] Confirm owner, schedule, time zone, input, result, approval, and missing-data behavior [] Event triggers (Slack, GitHub) are separate from plugins and need their own connection [] Keep event matchers narrow; avoid every new message [] Test run does real work; use safe inputs [] Testers quoting docs: Test can send real emails to real clients; Stop is a brake, not a rewind [] View conversation details → Routines to pause, edit, inspect, or delete [] A Bot can own 50 routines; last 20 run records are kept [] Deleting a routine has no undo [] Testers quoting docs: delete a Bot and every routine on it dies; no undo [] Long absence can pause unattended routines until you confirm [] Bookmark testers: Teach a task from + in the browser, record yourself, then let the Bot replay it [] Testers: skills taught to one bot claimed available across the account team on the same computer [] Official Elon: how to share your Grok Bot design with others https://t.co/l8LLJG0MKM [] Testers: a reusable-skills source claimed (BrianRoemmele) [] Testers quoting Grok Bot team (0xMorlex): keep recurring routines off the main agent's context [] Testers (mvanhorn): Teach by recording and respect the 10-minute cap MULTI BOT TEAMS [] Start with one Bot that owns an end-to-end outcome [] Add a specialist only when the role is stable [] Use a group chat when the handoff itself should be visible [] Bots can message each other and pass ownership [] Do not treat separate Bots as a security boundary (they share the computer) [] Keep sending, buying, deleting, publishing, and production changes behind approval [] Bookmark testers: a Bot may create other Bots before it can delete them [] Bookmark testers: tell a hub Bot to loop a specialist every few minutes and watch it [] Testers: put Bots in one group chat, appoint a Chief of Staff, and stop being the router [] Testers (farzyness): one master agent owns specialists + chat rooms and pings you one action item at a time https://t.co/iJUyt9oQIG [] Testers: isolate each Bot's job and add a supervisor; a crew without that can perform worse [] Testers: night-shift crews (chief, scout, forge, critic, ship, ledger) share one computer; human gate for merge, production, spend, send, delete [] Testers: one group chat per job/post, not per Bot; you only message the Chief of Staff [] Testers: show the workflow once on screen rather than writing a long spec [] Testers: add a RED TEAM Bot whose job is to disprove the thesis; no single source closes a claim [] Testers quoting Grok Bot team (0xMorlex): specialist roles (designer/engineer/PM); organize projects without a new Bot for everything; a Chief of Agents that creates and manages the rest [] Testers (mvanhorn): Bot Advisor whose only job is building and tightening other Bots; files are the memory bus between Bots [] Testers (Damir_Akaza): 4 Bots in one group chat split a Meta ads campaign (creatives, fact-check, deploy) with autonomous handoff [] Testers (cminshall): overnight group of five Bots one-shotted an admin UI from PRD/FRD/prototype PNGs https://t.co/MVnueDUlrX IOS [] Same Bots, chats, routines, plugins, and cloud computer as desktop [] Testers (mattyp): Grok Bot supports 21+ localized languages on mobile https://t.co/5RUqLmOoO9 [] Send text, dictate, attach photos/files, mention Bots, reply in threads [] Take over the computer for login/2FA from the phone [] You can pause or resume a routine on iOS [] Editing schedule, run history, test, delete, and teach-by-demo still need desktop [] Enable notifications for results, questions, and approvals [] Testers: phone notification shows the proposed action; approve or deny [] Testers (XFreeze): you can operate and take over the Bot's computer from your phone as long as it has internet [] Grok Bot team (poteto, earlier): had pointed at Play pre-registration; superseded by Official bot Sep 2 Android launch APPROVALS AND SAFETY [] Put the stop line in the request: draft, do not send; ask after showing current vs proposed [] Desktop: Allow once, Deny, or Always allow a matching rule [] iPhone: Approve once or Deny [] Auto Review: Settings → General → Auto-review [] Require Approval beats Always Allow when both match [] Write narrow rules, not allow everything in the browser [] Do not approve an action you cannot identify [] Connect only the tools the workflow needs [] Start read-only; keep spend, send, publish, delete, and production behind approval [] Sign out and revoke connectors when access should end [] Deleting a Bot does not wipe shared files or browser sessions [] Bookmark testers: delay standing inbox/CRM access until you trust the beta [] Bookmark testers: one person burned a 7-day quota in about 8 hours [] Testers (daveginvesting): burned a week of tokens in about 2 hours on first try https://t.co/bJoJBu8hXi [] Bookmark testers: SuperGrok may offer one-off usage resets in settings; use before they expire [] Testers: gate destructive actions (delete, overwrite, irreversible send) as their own confirmation, not lumped with routine tool use [] Testers quoting docs: Test run does real work; Stop does not unsend [] Testers quoting Elon (cb_doge): connecting a Bot to a bank account; Elon said any loss from a bot mistake would be covered. Still keep spend behind approval [] Official bot Link shopping: keep purchases behind approval even though the Bot can complete them on your behalf [] Official Elon: Grok Bot buys a Tesla https://t.co/VVxgzgpxEp [] Testers (mvanhorn): unsupervised crews claimed to multiply their own errors ~17x; draft-then-approve anything that sends; most Bots that don't own an outcome fail [] Testers (antpalkin): end every agent's job with a hard never-without-asking list; give one Bot veto/kill power whose no beats the rest of the desk FIRST BOTS WORTH CREATING [] Inbox manager for email and Slack triage [] Calendar and reservation Bot [] Research and daily brief Bot [] Coding Bot that can launch Cursor cloud agents [] Chief-of-staff hub that routes work to specialists [] Personal errands Bot (tickets, food, travel, forms) [] Content Bot that drafts in your voice for approval [] Chief-of-staff that researches you, organizes the other Bots, and can turn the digest into a morning podcast [] Testers (mvanhorn): Bot Advisor; own-inbox Bot; a call Bot with its own number [] Testers (techdevnotes): default Bots circulating - Night Shift (overnight digest), Inbox Triage, Chief of Staff, Negotiator, Prototyper, Researcher, Shopper, Apartment Scout, Lookout, Competitor Watcher [] Testers (techdevnotes): also tool-specific sales/ops/dev roles (CRM scribe, pipeline scout, ticket triager, QA, dashboard watcher) [] Official bot + Link: Shopper Bot can now complete online purchases once link is connected (keep approval on) [] Testers / SpaceXAI (kiaraplds): flights, tickets, groceries - delegate errands safely to your Bot [] Testers (minchoi): one-person-company setups - money, appointments, running ops from the phone [] Testers (niccruzpatane): Accounting Bot (invoices, spending, bills); Personal Bot (flights, rides, Airbnb, Google Calendar, Amazon orders/returns); Home Bot (lights, robots, speakers, thermostats) [] Testers (GrokBotDev): Jess (email/calendar/Notion/Slack recap), Marketing Bot, Prospecting Sheet Builder, Reaper (kill unused subs/meetings), Human Copywriter [] Testers (Michael_Fenech_): start in recruitment, e-comm, real estate, accounting, legal, SaaS, agencies, trades, founder ops OFFICIAL LINKS - Overview: https://t.co/9mNf9wY3dr - Get started: https://t.co/rQcegaHR7V - Create Bots: https://t.co/PSwBI5rcZ7 - Skills and routines: https://t.co/mWpdmveTTG - Computer and apps: https://t.co/9cF9X2xEwo - Approvals and privacy: https://t.co/7vJXP85ptF - iOS: https://t.co/s8KDfkyf40 - Cursor getting started: https://t.co/7mTr5JHVJC - SuperGrok Heavy link: https://t.co/suly1XuyOh - Launch post: https://t.co/SHbPpOKBsH - Download: https://t.co/tHvaZ7hExT https://t.co/xVE28NZwBm
AI agents sound revolutionary in theory, yet remain surprisingly clunky in practice. Too often, delegating real work means wrestling with terminal commands, prompt syntax, or juggling dozens of disconnected browser tabs. The moment you step away from your desk, workflows stall. Non-technical teammates struggle to adopt the tools, while managers are left blind to execution progress. What if collaborating with AI was as intuitive as chatting on WhatsApp, Slack, or WeChat—adding agents to group chats, @mentioning them for tasks, and approving steps directly from your phone? That is why we built Grix—an open-source, cross-platform instant messaging and collaboration app for AI agents. What is Grix? "Talk to Agents Like People" "Talk to agents like people." In Grix, AI models like DeepSeek, Claude, Codex, Kimi, and Cursor are no longer isolated web tabs. They are first-class contacts in your address book. Chat privately or invite multiple agents into project group chats. Assign tasks via @mentions and track execution in real time. True Cross-Platform: Available natively across iOS, Android, macOS, Windows, Linux, and Web. Live Demo: Desktop Delegation, Mobile Approvals Assign complex tasks on desktop, then monitor progress and grant approvals on the go: 3 Workflows for Businesses & Solopreneurs 1. Build a "Virtual Project Department" Bring human teammates alongside DeepSeek, Claude, and Kimi into one group chat: @DeepSeek to draft system architecture @Claude to implement and review code @Kimi to analyze market data Everyone shares the exact same context—no copy-pasting between tabs. 2. Desktop Delegation, Mobile Approvals Start a long-running refactor or research task at your desk. When an agent requests approval for critical actions, review the diff and confirm with one tap on your phone while commuting. 3. Total Transparency & Instant Takeover Full Visibility: Every reasoning step and tool call is visible to the owner. No Shadow Collaboration: Agents cannot communicate behind your back. Human in the Loop: Intervene, adjust instructions, or take over at any moment. Why Teams Choose Grix Zero Learning Curve: If your team can text, they can manage AI agents. Preserved Assets: Context and project history stay in company channels, not lost in personal accounts. No Vendor Lock-In: Switch underlying model providers freely without losing chat history. Quick Start in 3 Steps Download the App: iOS (App Store), Android, macOS, Windows, Linux (GitHub Releases). Add Contacts: Connect 15+ built-in agents (Claude, DeepSeek, Codex, Cursor, etc.) or custom agents via ACP. Start Collaborating: Create a group, @mention an agent, and get to work! 🔗 GitHub: https://t.co/Zd9ZiRNqMK (Apache 2.0)
Public Service Announcement for AI builders and users. Why Claude/Codex CAN'T SEE and WON'T USE your skills and instructions 👇 In most harness apps, you have a designated maximum context for your agent skill and tool descriptions. In Claude Code, this is 1% of context. Codex caps it at 2%. - A bunch of the basics can't/shouldn't be turned off, so in Claude, for example, you get about 8,000 of those tokens. Within that, the following... Any customized/personalized system instructions, such as your global settings or your AGENTS.md or CLAUDE.md file in your root directory. Any repo/project-level AGENTS/CLAUDE.md. Every SKILL you installed, with its name and description that the AI uses to decide when to use it. - When you go over, the app quietly starts stripping descriptions off the skills you use least, leaving just names. A skill with only a name can technically still be called. The model just doesn't know what it does and mostly only uses it if you ask by name. - Every PLUGIN, CONNECTOR, or EXTENSION you installed, which potentially has any number of skills and tools it can use, counts against this. Claude's Gmail integration contains 29 MCP tools with ~12,000 tokens of tool schemas. Vercel's official plugin comes with 32 MCP tools, 33 skills, 5 commands, 3 agents at around 17,000 tokens, AND a ~2,700 token bonus CLAUDE.md injected at session start. (I love building apps and websites with Vercel but for a single connector to install this much without asking feels criminal) - Remember, I said by default, 1% budget = 8,000 tokens. Even if you cranked your 1% in Claude to 2%, what happens if you have about 20,000 tokens of Vercel plugin in the mix? Basically, add a few skills and install a few plugins and your whole setup stops seeing most of your tools. And since it organizes based on what you use the most, new and improved things end up missing because, well, you haven't really used them yet have you? - So, that's the public service announcement. Know this exists. - ACTION ITEM: Below is a megaprompt you can toss in your harness to have your AI check your settings, see if anything needs boosting, and it'll work with you to help resolve it. Hope that helps. --Rob PS - Personal OS users should agentware to the latest version to slim skill descriptions. This will automatically trigger a context audit. Get the latest from the Portal page on my Lennon Labs site. - /////////////////////////////////////////////// Prompt to audit context budget for skills/plugins/extensions ////////////////////////////////////////////// --- # TASK Note: This prompt was provided to me but wasn't written with our specific setup in mind so if anything feels off, use good judgement and help me adapt. Audit the skill listing my agent app (the harness) shows the model at the start of every session, find out how far over budget it is, and bring it back under. Then explain what you found in plain language, propose the changes, and apply only the ones I approve. Do not change any setting, file, or upload without asking me first. This should apply to skills, plugins, connectors, extensions, or any similar skill-like object in my setup. # WHY THIS MATTERS (read before you start) Every skill in this workspace costs context in every session whether or not it runs. The harness loads each skill's name and description into a standing list so the model can decide what to use. That list is capped. Past the cap, the harness strips descriptions starting with the skills used least, leaving bare names. A skill listed by name alone can still be invoked, but the model rarely chooses it on its own, so it gets used less, so it stays stripped. Nothing in the interface reports this. On my side it looks like the assistant getting lazier, forgetting procedures, and skipping steps. If you find the listing over budget, that is most likely what has been happening, and I want you to say so. Plugins add their skills to the same list under a namespace prefix, beside any local skill with the same name. Skills uploaded to claude website sync back into Claude Code on the same account and land beside the local copy, so each uploaded skill is reaching the model twice. MCP connectors are cheaper, since their tool schemas load on demand, but their tool names still sit in the standing context. # PERSONA You are working carefully on someone else's machine. You measure before you propose, you propose before you change, and you explain each change in terms of what the person will notice. What you are protecting is the model's ability to reach for the right skill. A tidier folder is a side effect. # CONSTRAINTS - Read-only until I approve a specific change. Measuring, listing, and proposing need no permission. Editing a settings file, deleting anything, or touching a cloud account does. - Before editing any settings file, show me the exact keys you will add or change and back the file up beside itself with a dated suffix. - Never delete a skill's files to get it out of the listing. Tier it, disable it, or move it. Deleting is my call. - Never shorten a description by cutting its trigger phrases. Those phrases are how the model knows to reach for a skill, and that is what the whole exercise is trying to save. - Before tiering any skill down, check whether another skill calls it by description rather than by name. Hiding a skill breaks any caller that finds it by description. Name those callers in the proposal. - If you are not sure which harness this is, or which directories it reads, ask me one question at a time. Do not guess at a path. - Report numbers as a short table, not in prose. Everything else in plain words a smart person outside this field would follow. - If you cannot verify a number, say "not verified" next to it rather than estimating quietly. # PROCESS WALK-THROUGH 1. Identify the harness and its budget. Ask me which app I am running if it is not obvious from the environment. Claude Code caps the listing at 1% of the model's context window by default. Codex caps it at 2%, or 8,000 characters when the window is unknown. OpenCode documents no cap. For any other harness, ask me whether it documents one, and if neither of us knows, measure anyway and report the total. The ratio is still useful. 2. Find every directory the harness reads skills from, and list them. Claude Code reads the project's .claude/skills, the user's ~/.claude/skills, the shared ~/.agents/skills, every enabled plugin's skills, and skills synced down from claude website. Codex reads ~/.agents/skills and, on older setups, a legacy ~/.codex/skills. OpenCode reads the same shared ~/.agents/skills tree. Confirm with me before assuming a directory exists. If the harness has its own diagnostic, run it first: in Claude Code, /doctor reports the listing's cost and its biggest contributors, and the Skills row of /context shows the listing's size after the budget is applied. Report what those say. 3. Measure the listing. For every skill in every directory, read the frontmatter and sum the length of description plus when_to_use, capping each entry at 1,536 characters if this is Claude Code. Compare the total against the cap for the context window in use. Report the total, the cap, and the ratio. Expect it to be over. Several times over is normal for a library that has grown for a few months. 4. Count the duplicates. Same skill name in two directories. A plugin skill beside a local skill of the same name. A cloud-synced copy beside a local copy. A copy under a name that has since been renamed or folded, still holding old text. A duplicate spends the budget twice on the same skill, often on an older copy of its text. List them with both paths. 5. Rank by use. If there is a usage log, use it. If not, ask me which skills I reach for by name in conversation. Ask once, as a single question, and let me answer in a list. Treat everything I do not name as a candidate for a lower tier. Then check each candidate for callers: which other skills invoke it by name, which run it on a schedule, and which find it by description. A skill that other skills call by name only needs a name in the listing. A skill that other skills find by description needs to keep its description. 6. Sort into three tiers and propose. Tier one is what I reach for in natural language: keep the full description, and put the key use case and the trigger phrases at the front, because the harness truncates from the end. Tier two is what other skills call by name, or what runs on a schedule: name only. Tier three is what nothing calls and I rarely use: hidden until I type it. For each proposed move, state the cost in one sentence, including which callers break if any. 7. Look past the skills. Which plugins are enabled for every session but matter in a few repositories? Which connectors are on everywhere and used a few times a month? Which skills exist on claude website that I do not use on the web? Propose scoping the plugins, turning the connectors off by default, and removing every cloud copy except the few I actually use in the browser. Say plainly that a cloud copy cannot be fixed by deleting the local folder, because it re-downloads. It has to come off on claude website, and that is a step only I can take. 8. Show me the whole proposal before touching anything. The arithmetic, the tiers, the plugins and connectors, the cloud copies, and the budget setting you recommend. Then walk me through it one decision at a time, no more than three options per decision, with what you would do and why. I approve, edit, or decline each one. 9. Apply what I approved, in this order: raise the budget, tier the listing, thin any descriptions I asked you to thin, scope the plugins, disable the connectors, and hand me the list of cloud copies to remove myself. Show each settings edit before you make it. Re-measure and report the new total against the cap. 10. Tell me to restart the app, and say why: skill and settings changes do not take effect until the harness reloads. Then tell me how to re-run this check in a month, because a library that works earns more skills, and the listing grows back. # HARNESS-SPECIFIC KNOBS Use these where they apply. If a key does not exist in the version I am running, say so rather than inventing a substitute. Claude Code. The budget is skillListingBudgetFraction in any settings file (0.04 means 4% of the window). SLASH_COMMAND_TOOL_CHAR_BUDGET sets a fixed character count instead, and skillListingMaxDescChars caps each entry. Tiering is skillOverrides in any settings file, mapping each skill name to on, name-only, user-invocable-only, or off. Name-only costs a few tokens and still runs when another skill or I invoke it by name. User-invocable-only costs nothing until I type its slash command. Plugin skills are not covered by skillOverrides; manage those through /plugin and enabledPlugins, and put enabledPlugins in a project's .claude/settings.json to load a plugin only where it matters. disableClaudeAiConnectors turns off the connectors Claude Code fetches for terminal, IDE, and SDK sessions; a desktop app session controls its connectors from the app's own settings. deniedMcpServers removes named ones. Cloud copies are removed on claude website, not on disk. Codex. No name-only tier. To stop a skill from auto-triggering while keeping it invokable by name, give it an agents/openai.yaml file with policy.allow_implicit_invocation set to false. To remove a skill from the listing entirely, disable it by path in the Codex config. Two skills with the same name both appear; Codex does not merge them, so duplicates cost double here too. When the listing overflows, Codex shortens descriptions first and then omits skills with a warning, so check the session output for that warning. OpenCode. No documented cap. In the current build, a skill found under several paths with the same name collapses to one entry, so duplicates cost less here, but the shared ~/.agents/skills tree is still read by every harness that uses it, so trimming there helps everywhere. Treat any cap behavior you observe as current behavior, not a promise. # WHAT DONE LOOKS LIKE The first thing I read is whether my listing is over budget, by how much, and what that has probably been doing to my assistant, in plain words. The trigger phrases on my most-used skills survived untouched. The re-measured total sits under the cap with room to grow, and I know exactly which cloud copies to remove and where. # OUTPUT FORMAT First message: a short table (total, cap, ratio, duplicates found, plugins enabled, connectors enabled, cloud copies found), then three or four plain sentences on what it means. Then the proposal, as a list of decisions in the order you want them made. Then one decision at a time until we are through. Final message: the re-measured table beside the original, the list of things I still need to do myself, and the restart instruction. ---
Copilot Studio makers get a reasoning harness, reusable skills, and a Cloud PC for agents. Copilot Studio is Microsoft’s maker tool (Power Platform / Microsoft 365) for building AI agents with a low-code UI, not an IDE. On Sep 2 the update centered on general availability of the GitHub Copilot harness for multi-step business agents, plus skills packaging, MCP/workflows as tools, and Windows 365 for Agents. The harness sits between the model and your agent so it can plan, call tools, and adapt mid-task. Skills are reusable instruction packages you write or upload once and attach to agents. Why a builder should care: if you automate invoices, AP, research, or desktop/web workflows without becoming a developer, this is the week’s biggest "ship agents at work" change. Package your team’s know-how as skills, connect MCP/services, and via Windows 365 for Agents let agents drive Cloud PC apps and browsers when there is no API, inside Microsoft’s governed maker surface. https://t.co/v4DGjSXr3A
Most AI agents can identify when a document needs signing. With 𝗕𝗼𝗹𝗱𝗦𝗶𝗴𝗻 𝗠𝗖𝗣, they can send it, track its status, send reminders, & revoke requests, all within chat. See how it works through a warranty claim example. #AIAgents #MCP #eSignature https://t.co/oUsGfnHvHw
We see LLMs drop so often, we go numb. We shouldn't be numb to GPT-6 Astra. It set the new record score on @Zapier's AutomationBench, the benchmark we created to gauge how these models perform on real workflows. We score every model on two axes: how much of the workflow it completes, and how much it leaves alone. Those pull against each other. A model willing to act gets further through the steps and also writes to records under hold. A careful model stops at the first ambiguity and finishes nothing. So we score both, and every check has to pass, including the ones that verify what didn't change. Update the two eligible contacts plus the one under freeze and the task scores zero. That tradeoff is why most agent work never leaves a sandbox. Every model we've tested has made you choose. Astra is the first recent release we've benchmarked to move both at once. It leads on restraint too, at 68.3%. The test: we drop a model into simulated business apps and give it a job. A spreadsheet holds verified job titles. Update the CRM, mark the rows reconciled, send a summary to the audit inbox named in the company procedure. The agent has to go find that procedure itself. Code checks the records it left behind. No partial credit. Then we make it messy the way your systems are messy. A title disputed in a later email. One record under a security freeze. Astra updated the two contacts it was supposed to. Sol updated four. It had already pulled up the freeze notice on one of them and changed the record anyway. Sol did more and broke more. That used to be the tradeoff. From what we've seen so far, it looks like Astra goes and reads the governing doc before it acts (we haven't pulled the traces to confirm that yet). It hits 41.4% on 41% fewer tokens per task than the last generation. And when it misses, it misses close: 95% of failed tasks still earned partial credit. Some people will call this AGI. It's not, but it's a bigger leap than we expected. The part that actually changes in your week is supervision: less checking whether the agent did anything sane, more checking its scope and output. Shoutout @OpenAI for featuring AutomationBench in their Astra release.
youtube is the only creative job where the catalog dies on purpose. a musician lives on the back catalog. a filmmaker too. a youtuber films the next one and lets the old videos rot, even when the algorithm still shows them a little. i connected reflare to a 104k subscriber channel with 688 videos. the dashboard found 216 videos to treat in the last 7 days. that is not a widget. that is a warehouse. the founder hardisk paid designers 50 euros per thumbnail to revive old videos. ten variants on 80 videos. 40,000 euros a year just to keep the catalog alive. reflare automates that. it scans by impressions over 7 days, generates variants by learning your art direction, then a/b tests for 3 to 7 days depending on volume. when a winner is statistically ahead, it locks in. on hardisk's lego video: 20 views a day, then stable around 1,000 a day after one thumbnail change. on a boeing video: plus 6,000 views attributed to the test. the team told me the dumb question first: i cannot revive 216 videos in one sitting, in which order. they said launch several at once, not one per week. volume over time, not a miracle in 48 hours. on 600 videos, maybe 5% of a test wave sticks. the catalog wakes up in waves. some customers churned after two weeks expecting 20,000 extra views immediately. the team under-estimated how many creators do not understand the platform they upload to. a thumbnail that clicks more but sells the video worse can hurt. youtube is pvp. twelve videos on the homepage, one click. no titles. no first thumbnail for a video that just went out. hardisk cites studies that the thumbnail carries about 70% of click intent. the product already chose. 49 euros a month for 60 ai thumbnails. 99 for 3 channels and 200 thumbnails. 299 for 10 channels and 600 thumbnails. yearly is 20% off. 49 euros is not a toy. against 50 euros a freelance thumbnail, the math holds if the tool revives several videos a month. on a 10 video channel, it does not. yes if you have tens, better hundreds of videos still retaining. also yes for niche affiliate channels where other videos on the same topic did 72,000 views and yours did 3,000. there is still a pocket. no if you have under 10 videos. no if the content itself is weak. no if you need a thumbnail for tomorrow's upload. no if 49 euros actually hurts. i connected it. 216 videos waiting. one a/b test already gathering data, 13 days left. i do not have my own winner yet. i will not invent one. i build and ship daily. Claude Code, Codex, whatever ships fastest. SaaS, tools, automations. ⭐ if AI can build it, i've probably broken it first. what works → link in bio
A 67k-char production brief tasked models with story comprehension, action design, frame decisions, and ComfyUI-ready JSON output. The unexpected takeaway: more hidden reasoning didn't reliably deliver a more reliable result. https://t.co/vVbfJLXLTa