The first answer from an AI coding assistant is a draft, not a solution. Here is how I learned to drive AI toward well-built code, and why it rarely gets there on the first try.


Introduction

Over the past six months I have built three side projects with AI coding assistants at my side: LLM-Sim, YaAICV, and PocketKid. All three are written in Python, and all three were built faster than I could have built them alone.

But "faster" hides an important detail. The speed did not come from accepting what the AI produced. It came from learning how to challenge what the AI produced.

The single most important lesson of these six months fits in one sentence:

Never accept the first answer as the best answer.

This article explains what worked, what didn't, and the question I kept asking myself along the way: why doesn't the AI write the best code from the beginning?


What Worked

Scaffolding and boilerplate

Flask routes, SQLite schemas, Jinja2 templates, service worker setup for PocketKid's PWA: this is where AI assistants shine. Work that used to take an evening now takes minutes. The patterns are well known, and the AI has seen thousands of them.

Exploring unfamiliar territory

When I built the server-side PDF rendering for YaAICV, I didn't know the trade-offs between rendering libraries. Asking the AI to compare approaches, with pros and cons, worked as a fast and decent first map of the territory. Not a final decision, but a starting point.

Explaining code back to me

One of the most underrated uses: pasting a function and asking "what could go wrong here?" The AI is often a better reviewer than author, a point I will come back to.

Tests

Asking for unit tests, especially edge-case tests, consistently paid off. Tests became my main tool for verifying the AI's own code.


What Didn't Work

Trusting code that "looks right"

AI-generated code is fluent. It has sensible names, tidy structure, and confident comments. That fluency is exactly what makes it dangerous: code that looks professional is not necessarily code that is correct.

Large, one-shot requests

"Build me the approval workflow for PocketKid" produced something that ran, but it mixed concerns, skipped validation, and made assumptions I had never stated. Big prompts produce big, average answers.

Long sessions without resets

After many iterations in the same conversation, quality drifted. Earlier decisions were forgotten or contradicted. Starting a fresh session with a clean summary of the current state worked better than pushing on.

Outdated APIs

More than once, the suggested code used a deprecated library method or an older configuration style. It ran, with warnings, or failed in subtle ways.


A Concrete Example: The Wallet Update

PocketKid has a simple rule: when a parent approves a completed challenge, the child's wallet is credited. I asked the AI for the approval logic. The first answer looked like this:

def approve_completion(completion_id):
    completion = Completion.query.get(completion_id)
    child = Child.query.get(completion.child_id)

    completion.status = "approved"
    child.balance += completion.reward
    db.session.commit()

It's clean and readable, and it works in a demo. It's also wrong in at least three ways:

  1. No state check. Approving the same completion twice (a double tap on a phone, a browser retry) credits the wallet twice.
  2. Read-modify-write on the balance. Two approvals arriving at the same time can overwrite each other.
  3. No missing-record handling. An invalid ID crashes with an AttributeError instead of returning a clear error.

As a bonus, Query.get() is a legacy API in SQLAlchemy 2.x. The code runs, but it reflects older examples.

After I pushed back with "review this for concurrency and idempotency issues", the second iteration was much better:

from flask import abort
from sqlalchemy import update

def approve_completion(completion_id):
    completion = db.session.get(Completion, completion_id)
    if completion is None:
        abort(404)

    # Only a pending completion can be approved: makes the operation idempotent
    result = db.session.execute(
        update(Completion)
        .where(Completion.id == completion_id, Completion.status == "pending")
        .values(status="approved")
    )
    if result.rowcount == 0:
        db.session.rollback()
        abort(409)  # already approved or rejected

    # Atomic increment at database level, no read-modify-write in Python
    db.session.execute(
        update(Child)
        .where(Child.id == completion.child_id)
        .values(balance=Child.balance + completion.reward)
    )
    db.session.commit()

The AI knew how to write the second version all along. It only wrote it once I asked the right question.


Why AI Doesn't Write the Best Code from the Beginning

This is the core question, and once you understand the answer, it changes how you work with these tools.

1. It predicts likely code, not optimal code

A language model generates the most probable continuation of your prompt, one token at a time. I explored this in depth with LLM-Sim. Most of the code the model learned from is tutorials, examples, and quick answers. The most probable code is therefore typical code: the version that appears most often, not the version that is most robust.

The naive wallet update above is exactly what a tutorial would show. That is why it came first.

2. It doesn't know your context

The AI didn't know that PocketKid runs on a Raspberry Pi, that SQLite has specific concurrency behavior, or that kids double-tap buttons. "Best" code is always best for a context: scale, hardware, security model, team, maintenance horizon. Without that context, the model fills the gaps with generic assumptions.

3. "Best" is a trade-off, and the prompt rarely defines it

Readable or fast? Minimal or extensible? Strict or forgiving? Every design choice is a trade-off. If the prompt doesn't state the priorities, the model optimizes for the implicit goal of most prompts: something that works and looks complete.

4. Generation commits early

Because output is produced token by token, the model commits to an approach in its first lines and then stays consistent with it. It doesn't naturally step back halfway through and say, "this design is flawed, let me restart." A human engineer sketches, discards, and redesigns. A single generation pass doesn't, unless you ask for it.

5. It has no feedback loop by default

Unless it is running in an agentic setup that executes code and tests, the model never sees its code run. It cannot observe the race condition, the slow query, or the failing edge case. It produces code without the feedback that makes engineers improve their code.

6. It tends to agree with you

Assistants are tuned to be helpful, which often means following your framing. If your prompt suggests a weak approach, the model will usually implement it well instead of questioning it.

7. Its knowledge has a date

Training data has a cutoff. Libraries evolve, APIs get deprecated, best practices change. The model's sense of "current" is always somewhat behind.

The key insight: the model's capability is usually much higher than its first answer. Asking for a review, giving context, and stating constraints don't teach the model anything new. They pull out knowledge it already has.


My Workflow: Driving AI Toward Better Code

Over time, this became a repeatable loop:

[ Specify ]  context, constraints, priorities
      |
      v
[ Generate ] small, focused unit of code
      |
      v
[ Critique ] ask the AI to review its own output
      |
      v
[ Verify ]   tests, linters, my own reading
      |
      v
[ Refine ]   iterate, or reset the session

Specify before you generate

Instead of "write the approval function", I write:

Write a Flask function to approve a challenge completion.
Context: SQLite, SQLAlchemy 2.x, may run on a Raspberry Pi.
Requirements: idempotent, safe under concurrent requests,
clear HTTP errors for invalid states. Prefer readability over cleverness.

Ask for a critique, not just code

Some of the prompts I use most:

  • "Review this code as a senior engineer. List the five most serious weaknesses."
  • "What happens if this function is called twice at the same time?"
  • "What inputs would break this?"
  • "Propose two alternative designs and compare their trade-offs."

Verify like it's someone else's code

Because it is. My checklist before accepting any AI-generated code:

  • Do I understand every line? If not, it doesn't go in.
  • Are errors and edge cases handled (empty input, missing records, duplicates)?
  • Is it safe? Input validation, SQL injection, secrets not hard-coded.
  • Is it efficient enough for my context? No N+1 queries, no unnecessary calls to paid LLM APIs.
  • Does it use current library APIs?
  • Is there a test that proves it works, including at least one failure case?

Keep units small

Small requests produce code I can actually review. A 20-line function gets a real review. A 400-line file gets a skim, and that is where bugs hide.


Lessons Learned

After six months, my view of AI coding assistants is clear:

  • They are excellent accelerators and mediocre decision-makers.
  • Their first answer is a draft, roughly the quality of a capable junior developer working with no context.
  • Their real value appears when you treat them as a collaborator to challenge, not an oracle to trust.
  • The responsibility for quality never moves to the tool. It stays with the engineer.

In a way, this mirrors what I wrote about AI in enterprise sales: AI changes how the work is done, but judgment, context, and accountability remain human.


Conclusion

AI coding assistants didn't make me a faster typist. They made me a faster reviewer, designer, and tester, as long as I refused to settle.

The best code doesn't come from the first answer. It comes from the right questions.

If you are building side projects with AI, try one thing this week: before accepting any generated function, ask the assistant to critique it. You will be surprised how much better the second answer is, and what the first one was hiding.