AI coding without being a developer: building tools in Cursor and Claude Code
Intermediate11 min readNo-code AI Tools

AI coding without being a developer: building tools in Cursor and Claude Code

AI coding tools can help non-developers prototype small internal tools. This guide explains a test-first workflow, its limits, and when an experienced engineer must take over.

What you should be able to do

AI coding tools can accelerate a small prototype, but generated code is only a candidate implementation. Define acceptance tests, inspect every change, and involve an experienced engineer before production or consequential use.

Saved only in this browser.
In this article

AI coding tools can help a non-developer turn a narrowly defined requirement into a prototype. Tools such as Cursor and Claude Code can propose files, run commands, and help investigate failures, but they cannot certify that the result is correct, secure, maintainable, or fit for production. The person using the tool remains responsible for understanding the data, permissions, tests, and consequences.

This guide describes a bounded learning workflow for non-developers: what is reasonable to prototype, which evidence to retain, and where experienced engineering review becomes mandatory.

What you can realistically build

A realistic list of what AI coding tools enable for non-developers in 2026:

Realistic:

  • Internal tools and dashboards (Streamlit, simple web apps).
  • Automation scripts (Python, Node).
  • Custom integrations between existing tools (Zapier alternatives, custom webhooks).
  • Data processing pipelines (cleaning CSV files, extracting from PDFs, summarising documents).
  • Small browser-based games or interactive demos.
  • Personal productivity tools (custom note-takers, todo apps with specific quirks).
  • Slack bots, Discord bots, Telegram bots.
  • Static websites and landing pages.

Needs experienced engineering review:

  • Production SaaS products, where identity, tenancy, payments, monitoring, backup, incident response, and dependency maintenance become part of the product.
  • Mobile apps. The toolchain is more complex; the “ship it” path is harder.
  • Anything requiring deep system understanding (concurrency, distributed systems, performance optimisation).

Not realistic (yet):

  • Critical infrastructure or safety-critical systems.
  • Financial systems or anything regulated where bugs have legal consequences.
  • Anything where the failure mode is “user data leaked” or “money lost.”

The safest starting point is a local, reversible tool that uses synthetic or non-sensitive data. Value is a hypothesis until representative users run acceptance tests and the team measures the result.

The tools

Two main options in 2026:

Cursor. An AI-enabled code editor whose documented Agent tools can search, edit files, and run terminal commands. Review its current Agent tools documentation and disable or approve capabilities according to the project risk.

Claude Code. Anthropic’s terminal coding agent can edit files and run commands within its configured permissions. Use the current Claude Code setup guide and review every requested permission; terminal access can affect more than the project directory.

Other options worth mentioning:

  • GitHub Copilot coding agent. GitHub documents a repository agent that can work on a task and open a pull request for review. Availability and policy depend on the plan and repository settings; check the current GitHub task guide.
  • Replit Agent. A hosted builder whose own collaboration guide recommends planning, context, review, testing, and checkpoints. Hosting and data-handling requirements still need separate evaluation.
  • Lovable, Bolt, v0. Web-based “describe what you want and we’ll build it” tools. Great for prototyping landing pages and simple apps. Less powerful for ongoing development.

Choose a tool from current documentation, supported languages, data-handling terms, repository controls, and a small trial. Product rankings and feature availability change too quickly to treat a universal winner as fact.

The mental model

Working with AI coding tools as a non-developer requires a slight reframe.

You are still participating in software development even if you are not typing most of the code. The model translates some intent into candidate code. Your job is to:

  1. Describe what you want specifically and concretely.
  2. Test that it does what you want.
  3. Notice when something is off and describe what.
  4. Keep the system simple so you can understand what you have.

Clear requirements help, but they do not replace engineering knowledge. You define observable outcomes, review the diff, test normal and failure paths, and ask an experienced engineer to review anything you cannot safely evaluate yourself.

The 80/20 of effective AI coding

A few principles that separate users who succeed from those who get stuck:

1. Build small, build often

A large one-shot request makes requirements, changes, and failures difficult to isolate. Break the work into observable increments and require a test or inspection result after each one.

The fix is to build incrementally. Start with the smallest useful version. Test it. Add the next feature. Test it. Add the next.

A first project might evolve through small, testable increments like these (the labels are stages, not time estimates):

  1. Stage 1: “Make a script that reads a synthetic CSV and prints the rows where the email domain is .ee.”
  2. Stage 2: “Now make it also filter by signup date. Take the date as a command-line argument, validate it, and add tests for invalid dates.”
  3. Stage 3: “Now make it output a clean Excel file instead of printing. Preserve the original file and test empty and malformed input.”
  4. Stage 4: “Only after the script passes: propose a local web interface where I can upload the CSV and download the result. Explain upload limits, file deletion, and threat boundaries before coding.”

Each stage should produce a verifiable artifact. Whether it becomes a useful tool depends on the tests, data, environment, and reviewer competence; there is no reliable universal completion time.

2. Test every step

Each time the AI changes something, test it. Run the code. Look at the output. Confirm it matches what you expected.

This sounds obvious. The temptation, when the AI says “I’ve updated the script,” is to trust it and move on. Don’t. Run it. The model sometimes reports a fix that did not land. The faster you catch this, the cheaper it is to fix.

A practical habit: after every meaningful AI change, run the code. If you don’t run it, you don’t know if it works.

3. Read the code (a little)

You don’t have to understand the code line by line. But you should at least glance at what changed. Often you’ll notice something obvious: “wait, you removed the date filter — that wasn’t supposed to change.”

Cursor and Claude Code make this easy — they show you diffs of what changed. Glance at them. The 30 seconds you spend reading often catches the “AI helpfully refactored the thing I wanted to keep” failure mode.

4. Use git, even alone

Git is version control. It lets you save snapshots of your project and roll back if something breaks. Cursor and Claude Code can use git for you — you just ask: “commit this with the message ‘add date filter’.”

The discipline:

  • After each meaningful change, commit.
  • When the AI breaks something in a way you can’t easily fix, ask: “roll back to the previous commit.”
  • For bigger changes, create a branch first (“create a new branch called ‘add-email-feature’ and work there”).

Git can return tracked files to a committed state. It does not automatically protect untracked files, uncommitted work, databases, external services, or secrets, so keep tested commits and separate backups for stateful data.

5. Use a single small project at a time

The “many tools, one project at a time” rule. Resist the urge to have five half-built projects. Pick one, finish it (or get it to a useful state), then move on.

This matters because each project has its own context — its files, its dependencies, its quirks. Switching projects breaks the AI’s understanding of what you’re working on. Stay focused.

A worked example: building a real tool

Let’s walk through a real first project. The goal: a tool that takes a folder of customer-call transcripts, extracts action items and decisions from each, and produces a weekly summary.

This is a useful learning exercise, but duration varies with environment setup, data quality, API changes, and debugging. Use synthetic transcripts until data-processing terms, retention, access, and any required consent are approved.

Step 1: Set up.

Install Cursor (cursor.com). Open it. Create a new folder for your project. Open it in Cursor.

Step 2: Describe what you want.

In the Cursor chat, type:

I want to build a small tool. The input is a folder of .txt files (one per customer call transcript). The output is a Markdown file summarising decisions and action items across all the calls in the folder, organised by week.

Use Python and one model provider that our organisation has approved. Keep it simple — single script, no framework. Do not send real transcripts yet.

Walk me through the design first before writing any code.

Review the proposed plan. Confirm input boundaries, output schema, failure handling, data destination, cost limit, and acceptance tests before accepting code changes.

Step 3: Build incrementally.

Now let’s start with the smallest piece. Write a script that reads all .txt files from a folder and prints their names and file sizes.

Cursor writes the code. Run it. Confirm it works on a test folder of three sample transcripts.

Now add a step that reads each file’s content and prints the first 200 characters of each.

Run again. Confirm.

Now add the AI step. For each file, call the OpenAI API to extract decisions and action items. Use a structured prompt that asks for JSON output with keys “decisions” and “action_items.”

Run the tests against synthetic text. Configure the provider credential through its current official instructions and an approved secret store or environment injection; never paste it into chat, code, logs, or shell history.

Now aggregate the results from all files into a single summary document, organised by date (taken from the filename if possible).

Run. Confirm.

Now produce a Markdown file with the summary, written to the same folder, named “weekly_summary.md”.

Run. Confirm.

Each step ends with recorded test evidence. Watching code generation is not the same as understanding the implementation; use the resulting artefact only within the boundary you can independently review.

Step 4: Refine.

The extraction is missing implicit action items. When someone says “yeah let me look into that,” it should count as an action item with [implied owner: speaker]. Update the prompt.

Some transcripts have multiple speakers. The current prompt doesn’t track who said what. Update it to attribute decisions and actions to specific speakers when possible.

Add a “what’s surprising about this week” section to the summary, where the AI surfaces unusual patterns.

Each refinement is a small request. Each one is tested before moving on.

Step 5: Polish.

Make the output Markdown have proper headings, links to the source files for each action item, and a nice header with the date range.

Handle the case where the folder is empty or has no transcripts — don’t crash, produce a helpful message.

Add a small CLI: usage python summarise.py <folder>. Print help if no argument is given.

Step 6: Document.

Generate a README.md explaining what this script does, how to install dependencies, how to set up the API key, and how to run it.

You now have a documented prototype candidate. Before using real customer transcripts, add representative acceptance tests, dependency locking, failure handling, privacy and retention controls, access restrictions, and review by someone qualified to assess the code and deployment.

A prototype form is checked against fictional paper test records
AI-generated illustration of testing a small practical tool without using real personal data.

The traps

A few specific failure modes for non-developers using AI coding:

Trap 1: scope creep without testing. “Add this, also add that, also let’s add…” Without testing between additions, complexity compounds, and when it breaks you don’t know which addition broke it. Build small, test always.

Trap 2: trusting that the code works because the AI said so. The AI sometimes claims things work that don’t. Always run the code.

Trap 3: repeating guesses without new evidence. If successive changes do not improve a failing test, stop. Preserve the error and logs, return to the last tested commit if safe, reduce the reproduction, and ask a qualified reviewer rather than letting the model keep changing unrelated code.

Trap 4: deploying to production prematurely. A tool that works on your machine may have security issues, performance problems, or edge cases when other people use it. Be careful about what you deploy and to whom.

Trap 5: owning a system you cannot evaluate. A model’s explanation is not independent assurance. Learn the relevant runtime, dependencies, environment variables, API boundaries, logs, and tests—or keep the work as a disposable prototype under experienced supervision.

When to actually hire a developer

Some signals that what you’re trying to build has grown beyond AI-coding-as-a-non-developer:

  • You can’t describe the problem any longer; you have to describe code.
  • The tool has availability, concurrency, performance, backup, or incident-response requirements you cannot test.
  • You’re handling sensitive data (customer PII, financials, health) and the failure mode involves data loss or leak.
  • You need integration with complex enterprise systems.
  • You can no longer explain or independently test the relevant code and dependencies; line count alone is not a useful safety threshold.
  • You’re hitting bugs that the AI can’t fix, and they keep recurring.

At any of these thresholds, bring in an experienced engineer. A tested prototype can clarify requirements, but the engineer may keep, replace, or redesign it after reviewing the code, data flows, and operational risks.

This is a healthy pattern: non-developer builds prototype, developer productionises. Both kinds of work are real and they complement each other.

What this changes

For non-developers, AI coding tools change three things:

You can test ideas earlier. A bounded prototype can make an internal-tool request concrete before the team commits to production engineering. Do not promise a delivery date or time saving until representative tests support it.

You can prototype before specifying. Instead of writing a 10-page spec for a developer, you build a small working version yourself, show it around, and refine. Specification by prototype.

You become more useful to developers. When you bring a developer in to scale or harden something, you bring a working artefact, not a vague request. The communication is much clearer.

The constraint has loosened

AI coding tools can lower the cost of exploring a small software idea. The transferable skills are describing observable outcomes, testing the result, reviewing changes, protecting data, and knowing when your own review is insufficient.

Pick a low-risk internal problem. Build the smallest local version against synthetic data, record the acceptance tests, and stop before external access or real data until an experienced reviewer approves the next boundary.

AI assistance does not remove the need for software engineering; it changes which tasks a beginner can attempt with supervision. Treat generated code as untrusted input, keep the scope reversible, and let evidence—not the model’s confidence—decide whether the result works.

Read next

Continue through the same learning path with the next practical articles.