A decision paper can read as if the work is finished while its central question remains unresolved. Language models make it easier to organise options, find inconsistencies and present an argument. Those capabilities are useful. They do not establish whether the right problem has been chosen, which disadvantages are acceptable or who will answer for the consequences. Those are responsibilities of the people making the decision.
My position is that AI assistance belongs within an established approach to problem solving. The working method should keep the person's contribution to the judgment visible. Kupermann Decision Partner is an open set of instructions designed for that purpose. It structures a conversation with a language model around the work needed to reach and review a decision.
An older method for a new division of work
George Pólya's How to Solve It, first published in 1945, sets out four phases of mathematical problem solving. Understand the problem, devise a plan, carry out the plan and look back. This simple structure places execution within a wider process. Pólya, Princeton University Press
Those phases provide a starting point for the design here. The repository develops them into six steps, covering the problem, criteria, options, focused investigation, challenge and subsequent review. The process begins by identifying the decision, its constraints and what would count as an acceptable result. It closes with a record of the reasoning and a way to check what happens afterwards. Applying this structure to organisational decisions is a design choice in this project, not an application validated by Pólya.
For a consequential decision, recording an initial view can be useful when it gives the model specific assumptions to examine. An existing view can be used directly. If the person declines to formulate an additional view, the work proceeds. The purpose is a useful challenge, not a mandatory checkpoint. An initial judgment remains provisional, and better evidence can change it.
What the research supports
Risko and Gilbert review cognitive offloading, the use of external resources to reduce internal cognitive demands. Their article examines its conditions and consequences. It is neither an experiment with this method nor evidence that every instance of offloading damages thinking. Risko and Gilbert, 2016
Lee and colleagues surveyed knowledge workers about their use of generative AI. Greater confidence in AI was associated with less self-reported critical-thinking effort. These are reported experiences and correlations. The survey does not establish that AI caused a decline in cognitive ability. Lee et al., 2025
Buçinca and colleagues tested interventions requiring thought before accepting AI advice. Compared with interfaces offering simple explanations, these reduced overreliance on incorrect recommendations without a significant additional gain in overall task performance. Participants rated the systems as more complex. Buçinca et al., 2021
Bastani and colleagues studied different GPT-4 learning aids in a mathematics classroom field experiment. The design of the assistance affected subsequent performance without AI. Findings from education do not establish equivalent effects among experienced professionals or in management decisions. Bastani et al., 2025
I draw a design requirement from this evidence. Assistance should leave room for independent thought that can be inspected. That means traceable sources, visible assumptions and an opportunity to challenge the reasoning. The value of the additional effort depends on the decision at hand.
Responsibility remains explicit
The person determines the objective, constraints and acceptable consequences. They control which information can be shared and remain answerable to the people affected. The model helps organise the investigation, proposes alternatives and exposes missing information. Claims about the outside world need sources. Calculations need a visible derivation. A second persuasive explanation from the same model is not an independent check.
Kupermann Decision Partner keeps the initial position, options, evidence and objections within one working process. Open questions remain open until evidence resolves them or the decision owner explicitly accepts the uncertainty. The amount of work should reflect the stakes and the ease of reversing the decision. A routine choice needs less investigation than a system whose mistakes reach customers.
A pilot with an incomplete business case
Consider a fictional customer-service pilot. Eight experienced volunteers use an AI assistant on 80 routine tickets. They report saving six minutes per ticket, which a proposal then extrapolates to annual demand. The figure excludes four minutes of review. That review catches two incorrect statements about cancellation terms before they reach customers. These are invented assumptions for a worked example, not results from a client engagement.
The arithmetic changes immediately. Suppose annual demand is 24,000 tickets, of which 60 per cent are routine. Two minutes saved across 14,400 tickets would release 480 hours. At an assumed EUR 50 per hour, that represents EUR 24,000 of capacity value. Annual licences cost EUR 18,000, leaving EUR 6,000 before integration. A further EUR 12,000 in initial integration costs makes the first-year balance negative by EUR 6,000.
Even this corrected calculation does not establish a financial return. Released time becomes a cash saving only when avoidable expenditure actually disappears. The observations also come from experienced volunteers handling a small selection of routine work. They do not establish results for complex cases, other staff or normal operating conditions. Catching two errors shows that those checks mattered. It does not show that every error was caught.
The present evidence does not justify a full rollout. Focused validation is worth considering if its expected contribution to the decision justifies the effort and cost. It would measure complete handling time, including review and rework, against a comparable process without AI. Cases should represent the intended deployment, and quality checks should look for errors independently of the pilot users. Stop conditions and the economic threshold belong in the plan before the results arrive. If resolving the uncertainty is not worth its cost, declining another test is a reasonable outcome.
Use and limitations
The English-language repository provides the instructions, working templates and complete example. A suitable starting point is one bounded decision with a named owner. Existing information and judgments provide the starting point for the investigation. The final record states the recommendation, objections, remaining uncertainty and a concrete reason to revisit the decision.
The method takes effort. It requires source checking, attention and a willingness to abandon an attractive proposal. It cannot technically prevent false model answers or guarantee that the instructions will be followed in every conversation. Any benefit over competent, less structured work needs a comparison under equivalent conditions. An unchanged or worse result would matter too.
A useful method makes a decision easier to examine. A person must still answer for it.