Free GitHub Copilot Prompts for Unit Test Generation: Ideas That Actually Work
Copy-ready Copilot Chat prompts for unit tests in Python, TypeScript and Java, covering edge cases, mocks, regression tests and a review workflow that catches weak tests.
At a glance
- 10 prompts
- 10 min read
Jump to a prompt 10
- Pytest prompts for business logic
- Pytest prompts for business logic
- Jest tests with mocks
- Jest tests with mocks
- Characterisation and regression tests
- Characterisation and regression tests
- Edge-case and security-aware prompts
- Edge-case and security-aware prompts
- Worked example: weak to strong
- Worked example: weak to strong
A Copilot-generated test suite that passes on the first run is a warning sign, not a win. The usual output is a pile of happy-path tests that mirror the implementation, assert whatever the code currently returns, and give you a green coverage number without protecting anything. If you want Copilot prompts for unit test generation that produce tests you would actually merge, you need to give the chat the same things a careful reviewer would ask for: the behaviour contract, the edge cases that have bitten you, the framework and version, and a rule that tests must fail when the code is wrong.
This guide is for working developers. The prompts are written as Copilot Chat messages (or inline comments) and filled in for realistic code. Plans, models and features change, including what is free, so check GitHub's current Copilot plans and limits.
How Copilot fits unit testing
Copilot is good at boilerplate: arranging fixtures, parametrising cases, writing mock setups, and drafting a first pass for a function it can see. It is weaker at knowing your business rules, at deciding what behaviour is correct (it will often "bless" a bug), and at integration details that live outside the files in context.
Rules for every prompt below: attach the relevant files or select the code, name the framework and version, tell it the expected behaviour rather than letting it infer from the implementation, and run the tests yourself. Never accept a test you do not understand, and check that each test fails when you break the code on purpose (a quick mutation check by hand).
A prompt formula for tests
- Target: file, function, class.
- Stack: language, test framework, versions, style (pytest parametrize, Jest
describe). - Behaviour contract: inputs, outputs, errors, in plain rules.
- Cases to cover: normal, boundary, invalid, concurrency or time if relevant.
- Isolation: what to mock and what not to.
- Output rules: naming, file path, no snapshot of internals, no logic in tests.
Scenario 1: Lena, a Python dev with a pricing function (pytest)
Lena maintains pricing.py in a Django 4.2 shop on Python 3.11. calculate_total(items, coupon, tax_rate) applies a percentage coupon, rounds half-up to 2 decimals, and caps discounts at 50 percent.
Pytest prompts for business logic
Write pytest tests (pytest 8, Python 3.11) for calculate_total in pricing.py. Use this contract, not the current implementation, as the source of truth: items is a list of (unit_price: Decimal, qty: int); coupon is a percentage 0-100 or None; discount is capped at 50 percent; tax is applied after discount; the result is Decimal rounded half-up to 2 places; qty <= 0 or negative price raises ValueError; empty items returns Decimal("0.00"). Use pytest.mark.parametrize for the table of cases, include boundaries (coupon 0, 50, 51, 100), and a case where rounding changes the cent (e.g. 3 x 0.335). Test names must describe behaviour. No mocks needed. Put them in tests/test_pricing.py. If the implementation disagrees with the contract, tell me in a comment instead of changing the test.
Why it works: "contract, not the current implementation" stops Copilot from asserting existing bugs; the last sentence makes disagreements visible. What to expect / check: a parametrised table; failure is expected values computed with the same formula as the code. Recompute two by hand. Iterate: "Add property-style tests with hypothesis for: total is never negative and never exceeds the undiscounted total plus tax."
Review the tests you just wrote for calculate_total. List any test that would still pass if I (a) changed round-half-up to bankers rounding, (b) applied tax before discount, (c) removed the 50 percent cap. For each gap, write the missing test. Do not modify existing tests unless they assert the wrong thing.
Why it works: it turns Copilot into its own mutation checker. Check: actually make those three code changes and confirm the tests fail. Iterate: "Now list the three cheapest further mutations a reviewer might try and cover those."
Scenario 2: Kofi, TypeScript developer with an API client (Jest)
Kofi has src/api/orders.ts in a Node 20 project using TypeScript 5 and Jest 29. fetchOrders(customerId) calls httpClient.get, retries twice on 503, and maps the response to Order[].
Jest tests with mocks
Write Jest 29 tests in TypeScript for fetchOrders in src/api/orders.ts. Mock only httpClient.get using jest.mock; do not mock the mapper. Cases: (1) returns mapped Order[] for a 200 response with two orders; (2) returns [] for an empty list; (3) retries exactly twice on 503 then succeeds on the third call; (4) throws OrderServiceError after three consecutive 503s and httpClient.get was called three times; (5) does not retry on 400 or 404; (6) rejects when customerId is an empty string without calling httpClient. Use fake timers if retry delays use setTimeout. Use typed fixtures, no any. Name tests "should ... when ...". Assert on behaviour and call counts, not on private implementation details.
Why it works: each case is enumerated, so none is invented or skipped, and the mock boundary is explicit. Check: retry tests often hang without fake timers; look for jest.useFakeTimers and advanceTimersByTimeAsync. Iterate: "Add a case where the 503 response includes a Retry-After header and show how the test should assert the delay, if the code supports it. If it does not, say so."
Refactor the tests above to remove duplication using test.each for the status-code cases, and a small factory function makeOrderResponse(overrides) for fixtures. Keep each test readable on its own; do not hide assertions inside helpers. Show the full updated file.
Why it works: "do not hide assertions in helpers" prevents over-clever abstraction. Check: a failing test should still point clearly to the case. Iterate: "Show the output you would expect if the 503-then-success case failed, so I can judge the failure message."
Scenario 3: Mei, Java developer on a legacy service (JUnit 5)
Mei inherits InvoiceService (Java 17, Spring Boot 3) with no tests. She must refactor safely.
Characterisation and regression tests
I am about to refactor InvoiceService.generateInvoice in a Java 17 / Spring Boot 3 project and it has no tests. Write JUnit 5 + Mockito characterisation tests that capture its CURRENT behaviour, including any surprising behaviour, so I can detect changes during refactoring. Mock the repository and the clock (inject a fixed Clock). For each test, add a one-line comment saying whether the behaviour looks intentional or suspicious. Do not fix any bugs. List at the end the behaviours you were unsure about so I can ask the product owner.
Why it works: it separates "lock current behaviour" from "specify correct behaviour", which is the honest approach to legacy code. Check: suspicious items must be investigated, not silently locked in forever. Iterate: "Group the suspicious behaviours into a separate nested test class named KnownQuirks with a TODO per item."
Write a regression test for this bug: generateInvoice produces a duplicate line item when the order has two lines with the same SKU and different discounts. Reproduce it first with a failing test using the smallest possible fixture (two lines, SKU "A-100", discounts 0 and 10 percent). The test must fail on the current code. Then suggest the minimal fix, but keep it separate from the test. Include the assertion with an informative message.
Why it works: failing-first is the core of a good regression test. Check: run it against the buggy code and confirm it fails for the right reason. Iterate: "Add one neighbouring case: three lines, two sharing a SKU."
Edge-case and security-aware prompts
For this function (select code), list edge cases a thoughtful reviewer would test, grouped as: boundary values, invalid input, unicode and whitespace, null or missing values, large inputs, time or locale issues, and concurrency or ordering. For each give a concrete input and the expected result according to the docstring. Mark any where the docstring is silent as NEEDS DECISION. Do not write test code yet.
Why it works: separating design from code surfaces ambiguity. Check: the NEEDS DECISION list is a spec gap list. Iterate: "Now write tests for only the cases not marked NEEDS DECISION."
Write tests for parse_username(value: str) in auth/validators.py (Python 3.11) that reject hostile input: empty string, 10,000-character string, strings with null bytes, newlines, SQL meta-characters, unicode lookalikes, and leading or trailing whitespace. If the function uses a regular expression, add a test with a long near-matching input that would expose catastrophic backtracking, with a timeout using pytest-timeout. Tests must assert that invalid input raises ValueError and valid input is returned normalised. Do not execute anything against a real database.
Why it works: it names inputs rather than saying "test security". Check: the timeout test is the ReDoS guard; verify it fails against a deliberately bad regex. Iterate: "Add a case for each Unicode normalisation form of the same visible name."
Worked example: weak to strong
Prompt 1 (weak):
Write unit tests for this function.
Typical result: three or four tests with obvious inputs, expected values copied from running the function logic, no error cases, mocks of things that did not need mocking and 100 percent line coverage that proves little.
Prompt 2 (improved):
Fill in the blanks below, or click a highlighted word in the prompt.
Write pytest tests for the selected function. Contract: [PASTE 4-6 RULES]. Cover: normal, boundaries at [LIMITS], invalid types, and the bug we had last month: [DESCRIBE]. Use parametrize. Compute expected values by hand in comments, not by calling the function. Mock only [DEPENDENCY]. If a rule in the contract conflicts with the code, flag it.
Why it is better: the contract, the history of real bugs and the mock boundary give Copilot something to test against besides the code itself.
Workflow: from untested module to merged tests
- Ask for the edge-case list (above) and resolve the NEEDS DECISION items.
- Generate tests case by case for the agreed list.
- Run them; investigate any failures as potential real bugs first.
- Run the mutation-gap prompt, then break the code manually to confirm.
- Refactor for readability; keep assertions visible.
- Run in CI, check for flakiness by running the suite several times, then open the PR.
For related Copilot and testing guidance see Copilot prompts for database migration scripts and Bolt.new prompt examples for unit test generation.
Troubleshooting
| Problem | Likely cause | Fix | |—|—|—| | Tests mirror the implementation | No contract given | Provide rules and say they are the source of truth | | Over-mocking | No isolation rule | State what to mock and what stays real | | Flaky tests | Time, randomness or ordering | Inject clocks and seeds; avoid sleeps | | Wrong framework syntax | Version unspecified | State framework and version and show an existing test file | | Tests pass after you break code | Weak assertions | Use the mutation-gap prompt | | Huge unreadable test file | Everything in one prompt | Generate per function or per behaviour group |
Review checklist
- Every test has an assertion that could fail for a real reason.
- Expected values were not produced by the code under test.
- Edge cases from your bug history are covered.
- No secrets, real customer data or network or database calls.
- Tests run in isolation and in any order.
- Breaking the code on purpose makes at least one test fail.
- You understand every line you are merging.
FAQ
Is Copilot free for generating tests?
GitHub has offered different plans and free allowances over time. Check the current plans, limits and eligibility.
Should I trust the coverage number?
No. Coverage shows lines executed, not behaviour verified.
Can Copilot write integration tests?
It can draft them, but it cannot know your environment. Specify containers, fixtures and cleanup, and run them against test infrastructure only.
Should I paste proprietary code into chat?
Follow your organisation's policy and review GitHub's data-handling terms for your plan.
How do I stop it from changing my source code?
State "do not modify non-test files" in the prompt and review the diff.
Conclusion
Good Copilot tests start with a contract and end with a deliberate attempt to break them. Give it the rules, the bug history and the mock boundary, ask it to critique its own coverage, and keep the human decision about what correct behaviour is. Do that and you will get tests that guard your refactors rather than decorate your coverage report.