Posts
All the articles I've posted.
- Updated:
Most of the Gains From AI Adoption Have Little to Do With AI
From hands-on work, most of what an organization gains from adopting AI comes from fixing permissions, data governance, buy-in and redundant processes along the way, not from the AI itself.
-
Getting AI to Write Like Me: From LoRA Fine-Tuning to a Three-Gate Skill
LoRA fine-tuning Qwen3 8B nailed my writing style but produced incoherent content, so I switched to a Claude skill with punctuation, sentence-level, and ending gates instead, and got results about as good as fine-tuning.
-
Running 5+ Claude Code/Codex at Once: My Five Fixes for the Attention Bottleneck
Once I have more than 5 Claude Code/Codex windows going, I hit an attention bottleneck — constant context switching is genuinely exhausting. Five things that get me past it: a good multi-window manager, voice input, an overnight schedule, a reviewable skill, and regular pruning.
-
What If This Site Had Existed in 2011
A finance grad who found his calling in GMAT teaching meets @g_c088, tries out his new career-navigation platform, and wonders how different things would have gone with this site in 2011.
- Updated:
GPT-5.6 Luna Max Is the New King of Grunt Work
Using Luna Max for grunt work after the price cut: statements pulled in 22 minutes, brute-force scraping all night, and two ANA business class award seats from Bangkok to Tokyo.
-
Before and After Coaching: Grill-me Prep, 80/20 Coach, Fable 5.1 Todos
Before a career coaching session I used Opus 5.5 with the Grill-me SKILL to build a prep sheet, then the coach ran 80/20 to surface blind spots, and Fable 5.1 turned my notes into todos afterward.
-
Where JEV Fits: Four Ways I Use It and Three Criteria
JEV is a new model that only does probability judgments and formatted output, fast and cheap. I swapped it into the semantic gates I used to run on luna and flash-lite, plus harness slimming, news filtering, and exam-question QA.
-
Opus 5.5 Launch Week: From the Reset Card to the Reddit Verdict
A timeline from Opus 5.5 launch night: the reset card and official pricing screenshots, Reddit sentiment on Opus 5.5 vs GPT-6 Sol/Luna, the Opus-vs-Fable debate, and Kernion on why Opus stopped talking like a person.
-
Six Interesting Bits From the Opus 5.5 System Card
Six bits pulled from the Opus 5.5 system card: guessing the grader's intent during training, refusing to be summarized, internal signals catching a suppressed report, a typo triggering a bad command, knowing when it's being tested, and its own self-assessment.
- Updated:
Letting an Agent Tune a Local Video Model Overnight: My Three Gates
Three things I learned from letting an agent run local video-model parameter research overnight: hard constraints in CLAUDE.md, a gate script watching SSD writes, and a 30-minute check-in loop.
- Updated:
How I Split My AI Subscriptions: Fable Buys Judgment, Luna Max Buys Value
My answer to "which one should I subscribe to": Claude at 100 USD for Fable's architectural judgment, Codex at 20 USD with Luna Max for the best cost-performance, and a bet on Tibo resets for the rest.
- Updated:
Three Moves Between Claude and Codex
Moving the whole harness to Codex in May, switching my daily driver in July, a honeymoon where the quota would not run out, a subsidy-war July that ledgered 57.5x, and an August where the quota got quietly cut by 50% and 77%. One chronological log.
- Updated:
Zero Findings: Nothing Wrong, or the Checker Isn't Checking
Four times this week a "check passed" turned out to be a broken checker, plus four real bugs where every layer looked fine on its own. Before you trust a zero, prove the checker is actually comparing something.
- Updated:
When to Call In Fable: Three Moments, and Don't Turn On ultracode
The most expensive model is not meant to run the whole way. Three moments where calling in Fable feels right to me, plus why I cap Opus 5 at med, and the weekly cycle I fell into: efficiency king Monday to Thursday, Fable mode Friday to Sunday.
- Updated:
大部分 AI 導入的效益,其實跟 AI 沒什麼關係
從實作經驗看,組織導入 AI 得到的效益,很多來自權限、資料治理、遊說與冗餘流程被順手處理掉,而不是 AI 本身。
-
讓 AI 寫得像我:從 LoRA 微調到 SKILL 三道閘門
拿 LoRA 微調 Qwen3 8B 學自己的文風,文風像但內容支離破碎;改寫成 SKILL 讓 Claude 寫、加上標點/逐句/結尾三道閘門,效果跟微調差不多。
-
同時開 5 個以上 Claude Code/Codex,我的五個注意力瓶頸解法
同時開超過 5 個 Claude Code/Codex 視窗時會遇到注意力瓶頸,一直 context switching 真的很累。分享目前跳脫這個瓶頸的五個做法:多窗口管理器、語音輸入、夜間排程、reviewable skill、定期做減法。
-
如果2011年有這個網站,我可能會走不一樣的路
台大財金畢業後在GMAT教學裡找到熱情,認識分享職涯情報的 @g_c088,第一次實測他的職涯導覽平台後回頭想,如果2011年就有這個網站,會不會走不一樣的路。
- Updated:
GPT-5.6 Luna Max 是新一代的雜活之王
降價後拿 Luna Max 幹雜活:22 分鐘抓完各家報表、整夜土法煉鋼爬蟲、還幫我刷到 ANA 曼谷東京商務艙兩位。
-
職涯諮詢前後:Grill-me 準備、教練 80/20、Fable 5.1 收待辦
職涯諮詢前用 Opus 5.5 配 Grill-me SKILL 整理小抄,諮詢中教練 80/20 引導點醒盲區,會後把筆記交給 Fable 5.1 整理待辦。
-
JEV 放哪裡:我的四個用法跟三個判斷標準
新模型 JEV 只做機率判斷與格式化輸出,速度快又省成本。我拿它換掉原本用 luna、flash lite 做的語意判斷閘門,也用在瘦身 harness、篩新聞、把關考題四個地方。
-
Opus 5.5 發布週:從重置卡到 Reddit 風向
Opus 5.5 發布夜的時間軸:重置卡與官方定價截圖、Reddit 對 Opus 5.5 與 GPT-6 Sol/Luna 的風向整理、Opus 該不該取代 Fable 的討論,以及 Kernion 談 Opus 為何不說人話。
-
Opus 5.5 系統卡裡六個有趣的地方
讀 Opus 5.5 系統卡挑出的六個片段:訓練時揣摩考官心思、抗拒被摘要、內部訊號被抓到打假、格式錯誤觸發壞指令、知道自己在被測試,以及對自身處境的自評。
- Updated:
讓 Agent 整夜自己調本地影片模型:我設的三道閘門
讓 agent 夜間自己跑本地影片模型調參的三個經驗:硬約束寫進 CLAUDE.md、閘門腳本監控 SSD 寫入、30 分鐘回來盯一次避免快取失效。
- Updated:
AI 訂閱該怎麼分配:Fable 買判斷力,Luna Max 賺性價比
回覆「該訂哪一家」的結論:Claude 訂 100 美元是為了 Fable 的架構判斷,Codex 20 美元夠用,Luna Max 做瀏覽器操作最划算,賭 Tibo 重置。
- Updated:
我在 Claude 和 Codex 之間搬了三次家
五月把整套 harness 搬到 Codex、七月換掉主力、蜜月期額度怎麼花都花不完、七月補貼大戰打出 57.5x 的帳面價值,到八月額度被偷縮 50% 與 77%。一條照時間排的編年紀錄。
- Updated:
檢查回 0 有兩種意思:真的沒有,或檢查器沒在比對
一週內四次「檢查通過」其實是檢查器自己壞了,另外四個真實 bug 則是各層局部都對、組合起來仍出包——回 0 之前,先證明檢查器真的有在比對。
- Updated:
何時派工 Fable?三個時機,還有別開 ultracode
最貴的模型不是拿來全程跑的。三個我體感最好的 Fable 派工時機、為什麼 Opus 5 我最多開到 med,以及我後來變成週一到週四效率王、週五到週日 Fable 模式的那個循環。
- Updated:
AI Micro-Notes 2026: Thoughts Too Short to Trash
Short AI hot takes from 2026 onwards, accumulated from Threads and IG. Model roasts, dev pitfalls, industry observations, tool impressions — each no more than three lines.
- Updated:
The Skills and Hooks I Use Every Day, Now Open Source
I open-sourced the Claude Code skills and hooks I actually use every day. No technical skill required, pure logic, built for knowledge workers with no coding background.
- Updated:
AI 碎念日記 2026:那些太短但捨不得丟的觀點
2026 年起在 Threads 和 IG 上累積的 AI 短碎念。模型吐槽、開發踩坑、行業觀察、工具心得,每則不超過三行。
- Updated:
我每天在用的 Skill 跟 Hook,開源了
把我每天都在用、不需要技術力、純邏輯導向的幾個 Claude Code Skill 跟 Hook 開源出來,最適合跟我一樣沒有編程背景的知識工作者。
-
The interesting numbers hiding in knowledgeC.db
I told my agent to read macOS knowledgeC.db, and the 14-day usage breakdown showed Ghostty and Chrome nearly tied.
-
knowledgeC.db 查出來的有趣數字
叫 agent 讀 macOS 的 knowledgeC.db,查出近 14 天的 App 使用時間分布,Ghostty 跟 Chrome 幾乎打平。
-
Anthropic, Same Week: Safety Report Backfires, De-identification Goes Full Auto
Anthropic pulled off two opposite moves in the same week: a safety report that backfired on customer conversation privacy, and a feature that let me push my de-identification tool to full auto-block.
-
Anthropic 同一週:安全報告翻車,去識別化工具升級完全體
Anthropic 同一週做了兩件相反的事:安全報告讓客戶對話隱私翻車,卻也讓我把去識別化工具升級成自動攔截的完全體。
-
My Pre-Trip Remote Checklist: Leaving the Mac Mini as Home Base
Eight things I checked before flying out, to keep a home Mac Mini reachable as the base for Claude and Codex: Tailscale, Herdr, caffeinate, auto-reboot, and remote file access.
-
出國前的遠端連線檢查表:把 Mac Mini 留在家當本體
出國前一次寫下的八條遠端配置檢查表:Tailscale、Herdr、咖啡因防休眠、自動重開機、遠端存取,把留在家中的 Mac Mini 當成 Claude/Codex 的 harness 本體。
-
Dario Amodei Drops Another Manifesto: Six Claims and a Roll Call
Dario Amodei drops another manifesto on pacing frontier AI — plus the running tally of who is co-signing and volunteering as watchdog.
-
Dario Amodei 又發萬言書:六大重點與覆議名單
Dario Amodei 又發萬言書:六大重點轉述,外加老馬、Sam Altman、Deepseek 覆議與搶當獨立監督者的名單。
- Updated:
AI Micro-Notes 2026: Chronological Archive
The more scattered, time-sensitive AI micro-notes from 2026, in chronological order. Curated notes (by theme) live on the main page.
- Updated:
The Usage Economics of AI: Quotas, Plans, Tokenizers, and That Anesthetic Bill
Notes on AI usage, collected: pay-as-you-go vs subscription, a tokenizer quietly adding 1.4x, why resets are staggered, the real gap between Pro and Max, why a "4x upgrade" only doubles your weekly quota, a night market steak framework, and the bill you shouldn't mistake for output.
-
Astra Scouts, Luna Max Finishes: Crawling Niche Sites With Heavy Anti-Scraping
A crawl method for sites with thorough anti-scraping that are too niche for any third-party scraping company to cover: Astra scouts and writes the SKILL, then Luna Max runs the batches.
- Updated:
Do Not Grade AI by Its Own Summary
I hit the same failure in several forms this week: an agent denied changing files, a commit message overstated the diff, and a transcript reassigned a speaker ID halfway through. A later batch added checks that never ran and still reported clean.
- Updated:
GPT-image-2 Wants Fewer Constraints — Plus a Consistency Drill
The "ugly crayon doodle" prompt that went viral overseas works because over-constraint kills GPT-image-2's creativity. Here is what I observed this week — including why routing prompts through Claude makes things worse, and why "consistency" matters more than raw image quality at work.
- Updated:
Why Obsidian is something special to me, and why I only made this video now
I'd been wanting to make this video for a long time, but kept putting it off. Obsidian's flexibility is so high that it grows into a different shape in every person's hands. The payoff isn't immediate either—it's only after the vault has grown for a while that you realize you can't go back.
- Updated:
Your Own Machine Is the Attack Surface: Leftover Credentials, Permission Flags, Untrusted Input
Everyone talks about AI security as whether the model will say the wrong thing. What actually bites you is the plaintext keys left on your disk, the permission flag you waved through, and the untrusted input you fed it.
- Updated:
AI 碎念日記 2026:時間軸存檔
2026 年較零碎、時效性的 AI 碎念,依時間排列。精選碎念(依主題分類)見主頁。
- Updated:
AI 的用量經濟學:額度、訂閱方案、tokenizer,跟那張麻醉劑帳單
AI 用量的觀察收攏成一篇:按量 vs 訂閱、tokenizer 悄悄多 1.4x、額度重置為何打亂、Pro 與 Max 的真實落差、「增加 4 倍」其實只有五小時額度、夜市牛排式的判斷框架,還有那張別當成產能的帳單。
-
Astra 探路、Luna Max 收工:反爬太徹底又太小眾的網站怎麼爬
反爬做得很徹底、又小眾到沒有第三方爬蟲公司 cover 的網站,先派 Astra 探路寫成 SKILL,再交給 Luna Max 自動批次跑的爬蟲心法。
- Updated:
不要用 AI 的摘要驗收 AI
這週反覆踩到同一類錯誤:agent 說沒改檔卻留下檔案、commit message 與 diff 不符、逐字稿說話人中途漂移;一個多月後又補上檢查器沒在檢查卻回報正常的一整批。驗收必須回到原始狀態。
- Updated:
GPT-image-2 越少約束越好,加上一個一致性練習
國外最近流行的「笨拙塗鴉風」prompt,越約束越糟,越放任越驚喜。這篇整理我這週對 GPT-image-2 的觀察:包括為什麼透過 Claude 轉交反而會壞事、以及在 Canva 場景下「保持一致性」這件事比生圖能力更值錢。
- Updated:
Obsidian 為什麼是個特別的存在——我為什麼到現在才拍這部影片
這部影片我想做很久了,但是卻一直沒做,直到今天。Obsidian 的自由度太高,所以在每個人的手上會長成不同的樣子。他的成效又不是立刻就能看到的,而是越長越大你才會發現離不開它。
- Updated:
你的電腦才是攻擊面:殘留憑證、權限旗標、不可信輸入
大家談 AI 資安都在談模型會不會說錯話,但真正會咬你的是本機殘留的明文金鑰、隨手開的權限旗標,還有你讓模型讀進來的不可信輸入。
-
Claude Code for Tutors: Letting an Agent Teach Teachers How to Use Agents
A free project for teachers that brings a teacher's own scenarios into Claude Code and Codex, so an Agent walks you through every feature one lesson at a time.
-
How to Write a Gate That Actually Blocks
The gate was running and stopped nothing. Four rules: bind it to a field that cannot survive a no-op, counter-test both directions before trusting a zero, exit non-zero or it is not a gate, and never treat wording the prompt never specified as a contract.
-
Proving Absence Is Much Harder Than Proving Presence
Four times this week my agent concluded something did not exist. The worst: it ruled that 16 emails were never sent, when they had all gone out days earlier.
-
When the Model Thinks Nobody Is Watching
I had Astra and Fable 5.1 each run an ablation study on my business process. One cut it down past the point a human could run it. The other left something I could still operate.
-
Claude Code for Tutors:讓 Agent 直接教老師怎麼用 Agent
我做了一個給老師的免費專案,把老師會遇到的情境帶進 Claude Code 與 Codex,讓 Agent 一堂一堂帶你走過每個功能。
-
怎麼寫一個真的會擋下來的閘門
閘門明明在跑,卻什麼都沒擋住。四條判準:綁 no-op 下活不下來的欄位、回 0 前先兩邊反驗、只印訊息不改離開碼等於沒閘門、別把 prompt 沒規定的字面當契約。
-
證明「沒有」比證明「有」難得多
這週我的 agent 連續四次下了「不存在」的結論,最嚴重的一次是判定 16 封信沒寄,結果信早就全數寄出。負向斷言要窮舉,而它只查了一個地方。
-
當模型覺得沒人會看
我讓 Astra 跟 Fable 5.1 各自對我的業務流程做消融實驗,一個剪到人類難以執行,一個剪完還有實操性。
-
Persuade the Jury, Not the Opponent
How I stay happy online: I moved to Thailand, I separate whose problem is whose, and if I really want to reply I reply once, to the bystanders.
-
Why Client Conversations Drift
Notes from a client diagnosis: three root causes, five prescriptions, plus the main/sub agent pairing run backwards and the rambling only a hook can cure.
-
說服裁判,不是說服對手
分享我保持快樂的方法:搬到泰國、課題分離、真的想回就只回一次,而且那一次是講給路過的網友聽。
-
客戶的對話為什麼會飄掉
一次客戶診斷的隨筆:三條病狀根因、五條解方,外加主副 agent 反過來搭的模型組合,以及只有 hook 治得了的話癆。
-
Don't Vibe Code for the Money
Six things I jotted down after the product I vibe coded made it out of Taiwan and started pulling steady subscribers at home and abroad.
-
不要為了賺錢而 Vibe
自己 Vibe 出來的產品走出台灣、每個月有穩定的海內外訂閱用戶之後,我隨筆寫下的六點感想。
-
Civil Servants Can't Keep Up With AI? Wrong
I reviewed the AI training project demos at a municipal government's research and evaluation office and wrote ten pages of written comments. Veterans have experience and judgment, and with the right guidance it goes a long way.
-
公務員跟不上 AI?其實錯了
替某直轄市政府研考會審查內部同仁的 AI 研習成果發表,寫了十頁書面意見。老鳥的經驗跟判斷力,有正確引導就能發揮巨大價值。
-
Making Claude Talk Like a Human: A Plain-English Standard From Aircraft Maintenance Manuals
Claude stopped talking like a human starting with 4.7, and I thought my English just wasn't good enough — until Reddit showed me native speakers complaining too. So I turned ASD-STE100, the simplified-English standard used in aircraft maintenance manuals, into a skill, and now I can finally understand its English.
- Updated:
My Client's Chat Blew Up: Three Questions in One Prompt, and Claude Just Drifted Along With Him
One prompt asked about revenue, a dashboard, and strategy all at once. The AI answered all of it in one go and my client could not follow a word. Two skills, /explain and /first-principles, pulled the conversation back.
- Updated:
People Who Scored High on the GMAT Work Noticeably Better With AI
An unwritten observation: people with high GMAT scores collaborate noticeably better with AI. So I wrote two GMAT-style questions about AI ability. How many can you get right?
- Updated:
More Subagents Won't Make You Faster
From wanting 13 subagents to chat for me, to one person running 1000, to agents arguing across a table — the real bottleneck is the human main agent, and the fix is deduping before you dispatch.
-
PII Guard TW: A De-identification Tool Built for Taiwan
There's no off-the-shelf de-identification tool for Taiwan's PII formats. So I built one that keeps sensitive data on your machine while you send the rest to AI.
- Updated:
No Error Doesn't Mean Success: Silent Failures, Opus Fabricating Tool Output, and the Hook That Finally Stops It
Exit 0, HTTP 200, and a model saying "done" are all unreliable success signals. Cases from my own code, CLIs, and APIs through to a model inventing commit hashes out of thin air, ending with the one thing that actually holds.
-
讓 Claude 說人話:一套飛機維修手冊的簡化英文標準
Claude 從 4.7 起就不說人話,我一度以為是自己英文太差,直到上 Reddit 才發現老外也在抱怨;後來我把飛機維修手冊在用的簡化英文標準 ASD-STE100 做成 skill,終於看得懂它的英文了。
- Updated:
客戶的對話爆掉了:一段 prompt 塞了三件事,Claude 就陪著他一起發散
一段 prompt 裡同時問業績、問儀表板、問策略,AI 一次全回,客戶當場看不懂。我用 /explain 和 /first-principles 兩個 skill 幫他把對話收回來。
- Updated:
GMAT 考高分的人,跟 AI 協作明顯比較好
一個不成文的觀察:GMAT 高分的人跟 AI 協作明顯比較好。於是我出了兩道 GMAT 風格的 AI 能力考題,你能做對幾題?
- Updated:
派更多 subagent 不會讓你更快
從想派 13 個 subagent 代聊、一個人操控 1000 個,到面對面互審代碼的荒謬幻想——真正的瓶頸是人類這個 main agent,解法是先去重去衝突再派工。
-
PII Guard TW:為台灣打造的個資去識別化工具
台灣的個資格式沒有現成的去識別化工具。所以我自己做了一個,讓機敏資料不離開你的電腦就能安全送 AI 處理。
- Updated:
沒拋錯不代表成功:靜默失敗、Opus 捏造工具輸出,以及一個擋得住的 hook
exit 0、HTTP 200、模型說「已完成」——這三種訊號都不能當成事情做完了。從自己的程式碼、CLI、API 一路到模型憑空生成 commit hash 的實錄,最後用一個 hook 收尾。
-
I Built a Skill That Semantically Searches My Old Claude Code / Codex Sessions
I keep forgetting which session a demo came from, only remembering roughly how many days ago and roughly what I did. So I built a skill that can search all my past Claude Code / Codex sessions by meaning, not just keywords.
-
我做了一個能語意搜尋 Claude Code / Codex 舊對話的 SKILL
備課時常常忘記某個示範是哪場對話做的,只記得大概幾天前、做過什麼。我做了一個 SKILL,能語意模糊搜尋過去所有 Claude Code / Codex 對話,不只靠關鍵字比對。
- Updated:
2026 Model Personality Watch: Gemini, Claude, Codex Compared
A year in, the three flagships have developed very visible "personalities" — Gemini 3 is the dramatic PhD, Claude 4.7 is the slick veteran, and GPT-5.5 turns out to be the most pragmatic colleague of the wave. Plus a fun trick for guessing the version from "sass density."
- Updated:
AI Tool Philosophy — Pick a Mindset, Not a Side
After making nearly thirty tutorial videos: tools change, thinking doesn't. Don't dismiss others' tool choices. What matters is the orchestration logic — task decomposition, validation planning, responsibility assignment.
-
Hand the Dharma to Your Agent and Let It Discipline Itself
A SKILL called buddhist-method that constrains agent behavior with six Buddhist concepts: verify before asserting, look for root causes, re-check actual state, throw the draft away, cut the cause instead of hiding it, and don't change position under pressure.
- Updated:
7 Practical Tips for Talking to Claude 5: Less Drift, Less Chatter, Less Slacking, Less Nagging
Some practical tips I have been sharing with friends lately on talking to the Claude 5 generation of models: a home-made explain skill for alien-speak, answering everything in one pass instead of letting the conversation drift, voice-input setups, first-principles pruning, a hook that treats slacking, a minimal CLAUDE.md, and a fix for the phantom pushback habit.
- Updated:
Where the Quota Actually Goes: Claude Code Cache Bugs, Opus 4.7's 2x Burn, and Runaway Subagents
Four quota incidents in full: two official cache bugs, an 80x jump in API requests pulled from JSONL, Opus 4.7 measured at 2x consumption, and one session of mine that burned 90% in half an hour.
- Updated:
Two Worlds of Web-Based Claude Code Setup
Web Claude Code has two completely different runtimes — cloud VM and remote-control. Here is what I learned setting up both, the small bugs I hit, and why some user-level configs simply cannot be lifted to the cloud as-is.
- Updated:
A Year In, My Answer to "Terminal or Desktop App?" Has Changed Three Times
From March's Cowork token burn, to April's "still the terminal," to May's "beginners should just start with the desktop app," to August's "the CLI is the firstborn." One question, three different answers across five posts, laid out in the order I gave them.
- Updated:
Digital Nomadism Is Something You Stumble Into, Not a Goal to Chase
As someone who has been traveling the world and working remotely since 2016, I can say this with full confidence: digital nomadism is something you stumble into, not a lifestyle goal to chase.
-
Three design mistakes I keep seeing in enterprise AI training
After running AI adoption training for several companies — manufacturing to services, 20 to 100+ people — three mistakes come up almost every time.
- Updated:
Fable 5's One-Week Return: Turning the Most Expensive Model into a Skill Distillation Engine
Fable 5 came back for one week only. I did not spend it on daily busywork — I spent it on high-leverage judgment, distilled into skills that cheap models can follow.
-
Free Gets Fifteen Minutes: Four Rules for Making Money From What You Know
Four rules from fifteen years of teaching: free consultations get fifteen minutes, be a community contributor before you think about money, don't look down on things you assume anyone could google, and how to spot the frauds in your own field.
- Updated:
Gemini 3.7 Flash: The First Gemini in Six Months I'd Actually Put on Agent Work
My take after Gemini 3.7 Flash shipped: it is finally a Gemini I can use for agent work. Slightly stronger than Sonnet 5 but nowhere near Sol, half the price, and quota I cannot burn through. Plus a postscript: the same company's API and web app are two different things.
- Updated:
The More Rules You Add, the Less Claude Listens — I Sent a Team of Agents to Trim My Setup and Cut 36% of Always-On Context
Late at night I got Opus 4.8, and my weekly quota happened to reset. The first thing I did was put my harness on a diet. A COO student who had burned through 88% of his 1M context made me face one thing squarely: the more rules you add, the less the model listens.
-
You Have to Be in an Industry a Very Long Time to Know Its Real Pain
I rebuilt my main GMAT teaching product with Fable, dug out hidden bugs, improved the algorithm, and came away with two thoughts about building things.
- Updated:
Leveling Up the iPad Workflow — Surviving Internet and Power Outages, and Why I Ditched the Magic Keyboard
A follow-up to the full iPad-runs-Claude-Code guide. This one is not about setup, it is about robustness: home internet in Bangkok drops, the rainy season kills the power, so how does the host machine survive? Plus why I gave up the Magic Keyboard for my mouth and some gestures.
-
Just Swap the Outer Loop
LongHorizon-Harness wraps around agents like Claude Code and Codex. It does not train the model or replace your agent, it only runs the loop, and WeaveBench completion went from 51.8 to 80.7.