Hey everyone! I’ve been thinking a lot about how AI is changing what we should consider “working” software from a GTM perspective. My argument is that software is shifting from tools users operate to systems that produce outcomes autonomously. That changes how we should think about product, customer success, and go-to-market. I’d love to hear where you agree, where you disagree, or what you think I missed. Article: https://www.linkedin.com/pulse/how-ai-changing-what-working-software-means-danny-sheehan-mmtuc?utm_source=share&utm_medium=member_ios&utm_campaign=share_via For anyone who’d rather read it without leaving Slack, I’ve pasted the full article in the thread below.
How AI Is Changing What “Working” Software Means Al is one of the first enterprise software categories where vendors can't fully define what "it works" means before a customer ever logs in. The traditional enterprise software sales motion hasn't adapted to this reality. When you buy traditional software, expectations are straightforward: clicking button X produces Y every time. The behavior is documented, the feature either exists or it doesn't, and implementation is largely configuration, integrations, and user enablement. Al fundamentally changes that equation. An Al system's performance depends on the environment it operates in, the context it has access to, the quality of the underlying data, the prompts users provide, and the workflows it becomes part of. Vendors don't fully understand any of those variables until after the customer starts using the product. Despite these dependencies, Al is still sold in deterministic language:
"Generate meeting notes."
"Automate your outbound."
"Draft emails."
The expectation becomes that the product will simply work. When reality proves more nuanced, customer-facing teams absorb the gap between what was promised and what Al can reliably deliver in each customer's unique environment. Three things break the old model: 1. Probabilistic generation. Traditional software is deterministic-same inputs, same outputs. Generative Al doesn't work that way. The exact same prompt can produce multiple valid responses: one better structured, one more nuanced, one emphasizing entirely different details. None of them are "wrong."
The old question: "Did it work?
The new question: "Was this output useful enough, for this user, in this context?"
That is a fundamentally different success criterion than enterprise software has ever been built around. 2. Context dependence. An Al model doesn't operate in isolation. Its output depends on customer data, permissions, connected systems, conversation history, prompt quality, and workflow design. Two companies can buy the exact same Al product and have completely different experiences. The model isn't better or worse. The context is—and from the outside, you often can't tell which one is failing. 3. Evals measure the wrong room. Evaluations tell you how a model performs in controlled scenarios against predefined benchmarks. Customers care whether the product helps them do their job inside their own messy environment. A model can ace internal evals while failing a customer's expectations, because those expectations are shaped by company-specific context and subjective definitions of quality. Closing the GTM Gap The biggest challenge in enterprise Al isn't getting the model to work. It's getting the customer to agree that it's working. That requires more than a better model. It requires a Go-To-Market (GTM) motion built for probabilistic software. Here are three moves to make, in order: 1. Shift discovery from features to context. Stop selling Al as a magic button and start selling it as a system. Discovery can no longer be "what features do you need?" It has to become: "What data does this workflow rely on, and is it clean enough for an Al to read?" Then, tightly scope the initial rollout. Instead of a platform that "does everything," mutually identify 2-3 known workflows with sample inputs and anticipated failure cases. Create a constrained sandbox with clear boundaries— before the contract is signed. 2. Define "good enough" on Day 1. Customer Success cannot run traditional onboarding for probabilistic software. Set the expectation at kickoff that tuning prompts, building context, and refining outputs is not a bug—it's the reality of deploying Al. Then, define the acceptable threshold for each workflow. If you don't, every minor hallucination looks like a massive product failure, and support ends up debugging opinions instead of fixing actual product gaps. 3. Measure acceptance, and account for the verification tax. If a user generates a draft, deletes the whole thing, and rewrites it from scratch, the software successfully "executed"-and delivered zero value. Health scores have to be built on acceptance rates and usefulness, not API calls and clicks. The verification tax is the dangerous middle of Al adoption: the tool produces plausible work that still costs a human 20 minutes to clean, verify, and re-enter somewhere else. That's not automation. That's reformatted work, not removed work. l've watched health scores built on activity metrics flag the wrong accounts-and miss the ones actually at risk. Real telemetry ties to unglamorous outcomes: fewer missed commitments, cleaner CRM notes, faster routing. Not tokens generated. The Takeaway The vendors who win this era of software won't just build the smartest models. They'll be the ones whose GTM teams help organizations integrate Al into how they already work-instead of expecting customers to adapt to how the model works. How is your organization adjusting its expectations and success metrics for Al tools? Are you tracking activity, or true acceptance?
