Measure the real workload Compare prompts, source inputs, outputs, edits, cost, latency, failure modes, and human acceptance instead of evaluating impressive demos.
Find the hard-compute seam Score which decisions are stable enough for rules, templates, lookups, conventional code, or a purpose-built workflow.
Keep exceptions intelligent Route ambiguity and novel cases back to bounded AI or a person while routine work follows testable software with predictable cost.