![rw-book-cover](https://www.iandmacomber.com/thumbs/post-datastack.jpg) ## Metadata - Author: [[Ian Macomber]] - Full Title:: The Shape and Feel of the Post-Ai Data Stack - Category:: #🗞️Articles - URL:: https://www.iandmacomber.com/blog/post-ai-data-stack/ - Read date:: [[2026-09-06]] ## Highlights > Today the primary unit of SNL shows up passively on a feed: full skits on YouTube, subsets of skits on TikTok. People find podcasts through clips reviewed by LLMs and surfaced by algorithms, maybe never watching the original. SNL has reinvented itself to be broken up and reassembled by agents and algorithms. ([View Highlight](https://read.readwise.io/read/01m1v45bm3zd0n0mte4s9kpgk9)) ## New highlights added [[2026-09-16]] > the data equivalent of chopping SNL clips for TikTok. If you design data artifacts to be machine readable and ready for decomposition and reassembly by agents, you ensure your work shows up to drive decisions. ([View Highlight](https://read.readwise.io/read/01m2m7v7qw3jyfpynkep6mb1s9)) > Make your context and connectors as headless as possible, and do your best to ensure that whether someone is using a coworker, a coding agent, a BI tool, or a Slack bot, your data question gets the same answer. ([View Highlight](https://read.readwise.io/read/01m2m7z07sxmtnaewct1bxcyed)) > Every person can ask a slightly different question, look at a different slice of customer, use a different metric definition, end up with a different number, and draw a different conclusion. > **The scarce resource is consensus**. ([View Highlight](https://read.readwise.io/read/01m2m84ze473pf3ev8z5s777b1)) > You get there by **compounding context through evals**. > # pseudocode > results = ask_everywhere( > interfaces=["coworker", "coding_agent", "bi", "slack"], > question="What was net revenue retention last quarter?", > ) > assert same_metric_definition(results) > assert values_within_tolerance(results, relative=0.00001) > assert same_authorization_outcome(results) > assert required_evidence_used(results) ([View Highlight](https://read.readwise.io/read/01m2m86c51ewvv0gzskj5qdm7p)) ## New highlights added [[2026-09-17]] > As data questions get more personalized and dashboard building gets cheaper, the scarce resource becomes company-wide consensus. ([View Highlight](https://read.readwise.io/read/01m2qaqv219t17ed2es6gf5evc)) ## New highlights added [[2026-09-18]] > It’s trivial, cheap, and fast to aggregate 100m+ rows in a raw transaction table in Snowflake, it is non-trivial, expensive, and slow to to parse 100k raw Gong call transcripts via LLM. ([View Highlight](https://read.readwise.io/read/01m2scjrzpptcyzvpgvrjhkscq)) > **First:** companies without data teams and Snowflake/Fivetran bills (the vast majority) start their AI journey by plugging coworking tools and agents directly into their vendors. This skips the ETL and data modeling step. It lets individuals work fast, but it leads to reproducibility crises, where nothing is consistent across the org. This is why the first step of nearly every Applied AI solutions pitch is a forward-deployed team building a semantic layer that describes your business. ([View Highlight](https://read.readwise.io/read/01m2scnrec6bjhaa9a4rm2tdmh)) > every correction a data scientist makes is accumulated learning, which only compounds if you build the infrastructure to catch the error, trace the process, and distribute the fix to your entire company. ([View Highlight](https://read.readwise.io/read/01m2scpayfr0c0s4pe41v3rm4c)) > the best argument against connecting agentic vendors directly to your raw data without a semantic layer in between. It’s not just that the answers might be wrong, it’s that every correction goes toward teaching the vendor about your business instead of fixing your own stack. ([View Highlight](https://read.readwise.io/read/01m2scrc0gzkhmvh4exed8f8a9))