largely agree, though i'd put the reason somewhere else. the models don't write badly, they write fine. the tell isn't the prose, it's that the message has no cost behind it. a human written note implied someone spent twenty minutes on you, and that implication was most of what a cold email was ever buying. once the words are free the same words mean something different.
the research versus writing split is the right diagnosis but i think the boundary sits elsewhere. the two steps usually run separately and never share a view of what's worth mentioning. the model surfaces twelve facts, someone has to pick one, and the picking is the judgement nobody has automated, mostly because finding is measurable and choosing isn't.
the tool i'd actually want is one willing to come back with nothing when there's no reason to reach out. everything on the market is built to always produce a message, which guarantees a weak one every time the trigger isn't really there.