<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Experiment on Thomy Gölles</title><link>https://thomygoelles.com/tags/ai-experiment/</link><description>Recent content in AI Experiment on Thomy Gölles</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 26 Sep 2025 20:24:37 +0000</lastBuildDate><atom:link href="https://thomygoelles.com/tags/ai-experiment/index.xml" rel="self" type="application/rss+xml"/><item><title>Exploring the AI Experiment Rabbit Hole</title><link>https://thomygoelles.com/ai-experiment-rabbit-hole/</link><pubDate>Thu, 25 Sep 2025 22:00:00 +0000</pubDate><guid>https://thomygoelles.com/ai-experiment-rabbit-hole/</guid><description>&lt;h2 id="tldr"&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;I ran seven identical prompts (late Sep 2025) across ChatGPT (GPT-5), Microsoft 365 Copilot (Web), and Claude Sonnet 4 to see how they actually behave. ChatGPT and Copilot felt similar in structure, but ChatGPT was the most reliable on “what’s the latest now?” facts (e.g., Docker Desktop), while Copilot produced the most “CTO-ready” briefing—and was blocked exactly once by its Responsible AI layer on a trivial math question. Claude wrote in a friendlier, more guided style with a nice canvas, but it sometimes drifted on timelines/specs and once missed a table requirement. For workflow, exporting and navigation were clearly smoother in ChatGPT (and Claude), whereas Copilot’s UI (copy state, short title limit, odd rename error) slowed me down. In a bonus “write the post” challenge, both ChatGPT and Claude delivered usable drafts (Claude via an agentic setup), while Copilot fell short due to missing markdown project support.&lt;/p&gt;</description></item></channel></rss>