AI Development · Prompting · Practical AI

Build Better Apps with Fewer Tokens: Prompts, Models, and Cost

Learn how to write focused development prompts, choose the right AI model and reasoning effort, and read token costs without sacrificing a working, verified application.

Share

A long conversation can feel productive while the AI keeps rereading old decisions, exploring unrelated files, and rewriting code that already works. Token-efficient development starts with a clearer task: give the model enough evidence to act correctly, choose a model suited to the work, and define how you will know the change is finished.

This guide is for people building applications with coding assistants such as Codex. You will learn what consumes tokens, how to read the cost chart, and how to write prompts that reduce unnecessary work while keeping useful checks. The aim is a working, verified feature with less rework.

1. Understand what you are actually paying for

Tokens are units of text and other information processed by a model. A token is not the same as a word; code, punctuation, filenames, and different languages tokenize differently. Your typed prompt is only part of the input. Project instructions, attached files, chat history, tool definitions, search results, and command output can also enter the context.

  • Input tokens: material the model receives to understand the task.

  • Cached input tokens: an eligible unchanged prefix reused from an earlier request, charged at the applicable cached rate.

  • Output tokens: generated material, including billable reasoning tokens as well as the answer or code you can see.

OpenAI’s token and credit explanation describes the distinction between token usage, credit rates, and subscription allowances.

A five-line final answer does not prove that a task was cheap. The assistant may have read thousands of lines, made many tool calls, or used substantial internal reasoning before writing those five lines. Asking for a short summary helps control visible output; it does not set a hard limit on the total work.

2. Read the model cost chart correctly

Model cost and reasoning chart: API input, cached input and output rates, Codex credits, and calculated output costs from 1,000 to one million tokens.
Supplied model cost chart, dated October 7, 2026. Rates were checked against official OpenAI pricing and model documentation. The graph calculates output charges only; it does not predict task usage.

The chart’s lines answer a narrow question: how much would a given number of billed output tokens cost at each model’s listed rate? They do not tell you how many tokens a feature will require, how well a model will perform, or your total application-development bill.

At 100,000 billed output tokens, GPT-6 Luna costs $0.05, GPT-6.1 Sol costs $1.00, and GPT-6 Astra costs $5.00 under the chart’s Standard API rates. These are calculations at equal token counts, not measured costs for the same coding task.

Check current rates in OpenAI API pricing and the Codex credit rates below.

For sign-in-based usage, see Codex pricing and credits. Subscription usage depends on the plan and task; the credit table alone does not determine how many messages your allowance includes.

3. Choose the model for the next task

Treat model selection as a decision about the next piece of work. A small, explicit change has different needs from an ambiguous failure spread across multiple services. OpenAI currently recommends GPT-6.1 Sol for complex coding and agentic work, Luna for focused repeatable tasks, and Astra for the most demanding work.

Those roles come from OpenAI’s model guidance. Availability and controls vary by account, client, workspace settings, and rollout.

Development taskUseful starting pointWhen to reconsider
One component, explicit behavior, familiar patternGPT-6 LunaMove to Sol if the fix crosses layers or requires deeper diagnosis.
Feature across UI, validation, and persistenceGPT-6.1 SolMove to Astra if difficult tradeoffs or unresolved failures remain.
Ambiguous architecture problem or difficult debugging across systemsGPT-6 AstraReturn to a lower-cost model when the work becomes explicit and routine.

These starting points are practical recommendations, not a benchmark guarantee. Compare models on representative tasks from your own app. A cheaper model that needs repeated corrections can cost more overall; a stronger model that finishes accurately in fewer attempts can be the better choice.

Increase capability when you have evidence: the same failure returns, the explanation does not fit the logs, or the change requires relationships across several parts of the app. Hand the next model the exact failure and what has already been ruled out. Avoid asking it to rediscover the project.

4. Set reasoning effort deliberately

Model choice and reasoning effort are separate controls. Higher effort generally spends more tokens on internal analysis; it does not create a universal dollar multiplier per reasoning level. Supported settings depend on the model and client.

For GPT-6.1 Sol, start with the client’s default and adjust from the results. Official Codex guidance recommends High as a starting point for Luna and Light for Astra. A familiar, well-scoped task may need less effort than a difficult debugging problem. Increase effort when deeper analysis produces a useful improvement.

Max gives the selected model more time to reason about a single task. Ultra uses subagents for parallel work. Parallel work can reduce elapsed time, but it can also duplicate context and add coordination; it is not automatically token-efficient. Use it when the work has meaningful independent parts. Luna supports Max, but not Ultra.

See Codex model and reasoning controls for the current settings. Select the actual control in your client; writing “use High” in a prompt is not a reliable substitute for changing it.

5. Write a small, complete development brief

“Improve my app” leaves the assistant to choose the goal, inspect broad areas, and invent a stopping point. A useful prompt names the behavior, points to likely evidence, preserves important constraints, and defines a check. A slightly longer precise brief can save far more work than a short vague request.

PROMPT

Copy this: implement one clear feature

Implement this change in my application.

Goal: Disable the Save button while a profile save request is pending,
then restore it after success or failure.

Start here: src/components/ProfileForm.tsx and its existing tests.
Inspect the save handler and related code before changing behavior.
Expand to other files only when needed to complete this change.

Constraints:
- Reuse the existing request and button components.
- Preserve unrelated changes and the current API contract.
- Do not add a dependency for this task.

Acceptance:
- A pending request cannot be submitted twice.
- Success and failure both restore the button.
- Existing validation still works.

Run the relevant checks and report their results.
Final response: changed files, behavior, verification, and any remaining issue.
Do not repeat whole files in the response.

Replace the paths and acceptance checks with your real ones. The likely files are starting points, not permission to ignore a necessary dependency. “Smallest change” should mean the smallest complete fix that meets the behavior and passes the relevant checks.

Official reasoning prompting guidance recommends direct instructions and specific success criteria. Ask for evidence and a concise explanation; a transcript of internal reasoning is unnecessary.

6. Use discovery when you do not know the cause

When the failure is unclear, a short diagnosis can prevent speculative edits. Supply the exact error, reproduction steps, recent relevant change, and expected behavior. Share the useful portion of the log rather than an unrelated dump. Let the assistant request more evidence when needed.

PROMPT

Copy this: diagnose before editing

Diagnose this failure before changing files.

Expected: [specific behavior]
Actual: [specific behavior]
Reproduce: [short steps]
Exact error: [relevant error text]
Recent relevant change: [change, or unknown]
Start here: [file or module, if known]

Read the nearest implementation and relevant project instructions.
Search for the error or symbol and follow its call path.
Return the most likely cause, the evidence, and the smallest complete fix.
Ask only questions that would materially change the fix.
Do not modify files yet.

This is deliberately a read-only prompt. After you review the finding, authorize the proposed fix. For a routine change with a known solution, skip a separate diagnosis turn and use the implementation brief directly.

PROMPT

Copy this: interrupt a repeated failure

The last attempt still fails with:
[exact failing command and relevant output]

Already tried: [change and observed result]
Known working: [relevant check or behavior]

Check whether the previous assumption fits this evidence.
Gather new evidence before making another speculative edit.
If the cause is clear, make the smallest complete correction and rerun
the failing check. Summarize what changed in the diagnosis.
If essential information is missing, identify exactly what is needed.

Repeating “try again” gives the model little new information. A failed check is evidence: use it to change the diagnosis, narrow the search, or justify a stronger model.

7. Keep context useful and make reuse measurable

  • Point to relevant files instead of pasting the entire repository. Ask for targeted searches and bounded reads, with expansion when the evidence requires it.

  • Keep persistent project instructions concise. Put specialized instructions near the code or workflow that needs them.

  • Request the needed change and a brief explanation. Avoid requesting a full rewritten file when an edit or patch is sufficient.

  • Continue a useful conversation while its context still helps. When the goal changes or history becomes crowded, carry a compact handoff into a new chat.

For an API integration, place stable instructions and reusable reference material before task-specific content. Caching depends on a matching rendered prefix; related wording or an ongoing session does not guarantee a hit. Keep the model and shared prefix stable when reuse is useful, and check the usage data.

For GPT-5.6 and later models, API caching has a minimum cacheable prefix of 1,024 tokens and separate cache-write pricing. A write is charged at the cache-write rate instead of the ordinary input rate; it is not an additional charge on the same token. Codex credit billing has no separate cache-write charge. Keep context relevant and measure actual reuse before assuming a saving.

For cache requirements and accounting, see OpenAI’s prompt caching guide.

PROMPT

Copy this: create a compact handoff

Prepare a handoff for a new chat in at most 250 words.

Include:
- Current goal and accepted requirements.
- Relevant files and key decisions.
- Changes already made.
- Checks run and their observed results.
- Exact unresolved failure or next action.
- Constraints that the next assistant must preserve.

Distinguish facts from assumptions. Use file paths.
Do not repeat logs or whole code files. Do not omit an unresolved risk
just to meet the length target; flag if more detail is essential.

8. Calculate a request cost without double-counting

Here is a hypothetical GPT-6.1 Sol API request at Standard speed, within the chart’s short-context tier. It is an arithmetic example, not a measurement of a real development task.

Token categoryExample usageRate per 1MCalculated cost
Ordinary input10,000$2.00$0.020
Cached input reads20,000$0.10$0.002
Cache-write input4,000$2.50$0.010
Billed output, including reasoning5,000$10.00$0.050
Total token charges34,000 input + 5,000 outputStandard API rates$0.082

Calculate each category as tokens ÷ 1,000,000 × its rate, then add the charges. Input categories are disjoint: if total input includes cached reads and cache writes, subtract both to find ordinary input. Reasoning tokens are already part of billed output; do not add them again. Tool charges and other applicable fees are excluded from this example.

At the chart’s Codex credit rates, 5,000 GPT-6.1 Sol output tokens equal 1.25 output credits before input usage. That is a credit calculation, not a conversion of the API example into a subscription bill.

9. Verify the result and track cost per accepted change

Saving tokens by skipping a necessary check can create an expensive repair later. Define the verification that matches the change: a targeted regression test for behavior, type or build checks where relevant, and a rendered interaction check for UI work. Keep the checks proportionate, and report a blocked or untested case honestly.

PROMPT

Copy this: verify the accepted behavior

Verify the change against these acceptance criteria:
[paste the criteria]

Run the relevant checks for the changed behavior.
For a UI change, inspect the rendered screen and exercise the affected
interaction where the environment supports it.
Expand verification if a failure or dependency gives a concrete reason.

Report what passed, what failed, and what could not be verified.
Include the failing command or reproduction when there is a problem.
Do not claim completion based only on a successful build.

Compare several representative tasks using tokens, elapsed time, correction attempts, and the final acceptance result. Track cost per accepted change, including failed attempts and review. Equal token prices do not mean equal development outcomes.

API builders can inspect usage.input_tokens, cached_tokens and cache_write_tokens in input_tokens_details, usage.output_tokens, and reasoning_tokens in output_tokens_details. A prose request such as “use fewer than 2,000 tokens” is guidance, not a dependable hard budget.

In the Responses API, max_output_tokens caps generated output including reasoning. Too small a cap can produce an incomplete response, even before a visible answer. It does not cap input, tool fees, or an entire multi-request agent run. Handle incomplete status and measure real usage before tightening limits.

See the reasoning and output-limit documentation for usage fields and incomplete responses.

For your next feature, write one clear behavior, give the assistant the nearest useful evidence, choose a suitable model and effort, and state the check that proves it works. Keep the context and final report focused. Judge efficiency by how reliably you reach that accepted result.

Pricing and model guidance checked October 7, 2026. The supplied chart is included as a dated reference; confirm current rates and available controls before using it for a future budget.