·
8 min
Negative Testing: How to Derive Negative Test Cases Step by Step

Roman Kirchmeier - Autemos

Negative testing checks that a system rejects invalid input, a wrong sequence of actions, or use it was never designed for in a controlled way: with a clear error message and unchanged data. Missing error handling is expensive: in an OSDI study of five distributed data systems, 92% of catastrophic failures traced back to it. The examples below come from Swiss payments: IBAN, QR-IBAN, currency, and amount.
TL;DR: Negative testing uses a system in a way it isn't intended and expects a controlled rejection. Derive the cases from invalid equivalence partitions, boundaries, business rules, and error guessing. The German ISTQB syllabus 4.0.2 recommends testing one invalid partition per test case. In an OSDI study, 92% of catastrophic failures traced back to faulty error handling.

Figure 1: Positive and negative testing compared
What is negative testing according to ISTQB?
Negative testing is, in the official ISTQB glossary, “A test type in which a component or system is used in a way that it is not intended” (ISTQB Glossary, 2026). Listed synonyms: invalid testing, dirty testing. It checks the defenses: inputs, actions, or states the system should reject or ignore.
Many pages still quote older wording from an unofficial glossary copy. For audit documentation, cite glossary.istqb.org, which carries the definition in English and German.
The ISTQB glossary has no entry for positive testing (checked September 27, 2026). The term is industry usage: a positive test uses valid input to confirm the intended flow, the happy path. For the anatomy of a test case, see our guide to writing a test case.
How do positive and negative testing differ?
Positive tests confirm that valid input produces the expected result; negative ones confirm that invalid input is rejected in a controlled way. The ISTQB syllabus sets the order: positive test cases first, the negative ones after them (ISTQB CTFL Syllabus 4.0.1, 2024).
Aspect | Positive | Negative |
|---|---|---|
Guiding question | Does the system do what it should? | Does the system refuse what it shouldn't accept? |
Input | Valid values from valid partitions | Invalid values, invalid partitions, values past a boundary |
Expected result | Processing and correct output | Rejection, clear message, unchanged data |
ISTQB glossary | No entry | Entry, synonyms invalid testing, dirty testing |
Order in ATDD | First | After the positive test cases |
Payment example | CHF 250.00 to a valid IBAN | CHF -250.00 or currency USD |
Typical finding | Wrong calculation, missing step | Silent acceptance, crash, error message that leaks internals |
Invalid cases belong in the requirement itself. The syllabus (section 4.5.2) says acceptance criteria should describe both positive and negative scenarios. See acceptance criteria with positive and negative cases for how to write them into a user story.
Why is testing error handling worth the effort?
In an OSDI study of five distributed data systems, faulty error handling caused 92% of catastrophic failures. Yuan et al. analyzed 198 failures in Cassandra, HBase, HDFS, MapReduce, and Redis and found that “almost all (92%) of the catastrophic system failures are the result of incorrect handling of non-fatal errors explicitly signaled in software” (Yuan et al., OSDI, 2014).

Figure 2: Error handling as the cause of catastrophic failures in five distributed data systems (Yuan et al., 2014)
Many were cheap to find. In 58% of the catastrophic failures, simple testing of the error handling code would have caught the fault. 35% followed three trivial patterns (as of 2014): empty handlers or handlers that only log, overly general exception aborts, and code marked TODO or FIXME. 77% could be reproduced by a unit test. The study covers distributed data systems, not banking software.
Few teams test this code on purpose. In a survey of 154 developers, respondents said “in 70% of the organizations there are no specific tests for the exception handling code.” Only 27% had error-handling policies (Ebert, Castor, Serebrenik, JSS, 2015). These figures are self-reported.
A banking example: the missing hard block
In 2024 the UK Financial Conduct Authority fined Citigroup Global Markets Limited £27,766,200. A trader meant to sell a US$58m basket of equities, made an input error, and created a basket worth US$444bn. Controls blocked US$255bn, US$189bn reached an algorithm, and US$1.4bn of equities were sold (FCA, 2024).
The FCA found there was “no hard block that would have rejected this large erroneous basket of equities in its entirety.” This was a trading control failure, not one missing software test. It raises the core question: what happens to input that is formally possible and commercially implausible? A hard upper limit is an invalid partition you can test.
How do you derive negative test cases step by step?
The derivation takes five steps: find invalid partitions and boundaries, check business rules, guess likely errors, isolate each invalid partition, and fix the expected result up front. For 100% equivalence partitioning coverage, the CTFL syllabus counts “all identified partitions (including invalid partitions)” (ISTQB CTFL Syllabus 4.0.1, 2024).

Figure 3: Deriving negative test cases in five steps
Identify invalid partitions and boundaries. An invalid partition holds values the test object should ignore or reject, or for which no processing is defined. Pick one representative per partition and test values just past each boundary. The technique is covered in our article on equivalence partitioning and boundary value analysis.
Separate syntactic and semantic rules. The OWASP cheat sheet on input checks separates correct syntax of structured fields (date, currency symbol) from correct values in the business context (start date before end date, price within expected range) (OWASP, living document). Derive at least one violation per rule.
Guess likely errors. Error guessing uses known failure sources; the syllabus lists input errors such as correct input not accepted and parameters wrong or missing. Fault attacks, after Whittaker, provoke failures on purpose: an empty mandatory field, a double submit, an expired session, steps in the wrong order.
Test one invalid partition per test case. The German syllabus 4.0.2 added this rule in 2025: invalid equivalence partitions should not be combined in one test case, to avoid fault masking (ISTQB CTFL Lehrplan 4.0.2, 2025). The English 4.0.1 syllabus lacks that sentence. Keep every other field valid.
Fix the expected result before you run. Record error code, message text, marked field, and the data state after rejection. Without this test oracle, all you'll get is “the system reacted somehow.”
What does fault masking look like?
A test case sends amount -50.00 and currency USD. The system rejects the currency and stops checking, so nobody knows whether the amount check works. Two separate test cases answer both questions.
Invalid paths in longer flows, such as an aborted payment approval, belong in a dedicated test scenario with an exception path. The ISTQB Test Analyst syllabus defines the exception scenario with “(e.g., abnormal use or invalid input)” (ISTQB CTAL-TA 4.0, 2025).
Which invalid inputs must a Swiss QR-bill payment form reject?
A Swiss QR-bill payment form must reject, at minimum, a wrong IBAN length, wrong IBAN check digits, a mismatched QR reference, a wrong currency, and an invalid amount. A Swiss IBAN has 21 characters, for example CH93 0076 2011 6238 5295 7 (SIX IBAN, 2026). Each row holds exactly one invalid partition; all other fields are valid.

Figure 4: Checklist of invalid inputs for a Swiss QR-bill payment form
No. | Field | Invalid partition | Test value | Expected reaction |
|---|---|---|---|---|
1 | IBAN | 20 instead of 21 characters | CH93 0076 2011 6238 5295 | Rejected, length hint |
2 | IBAN | Wrong check digits (MOD 97-10) | CH94 0076 2011 6238 5295 7 | Rejected, invalid IBAN |
3 | QR reference | QR reference with a regular IBAN (IID 00762 is outside 30000–31999) | CH93 0076 2011 6238 5295 7 plus QR reference | Rejected, QR-IBAN required |
4 | QR reference | 26 instead of 27 digits | 26-digit number | Rejected, format error |
5 | Currency | Neither CHF nor EUR | USD | Rejected |
6 | Amount | Zero | 0.00 | Rejected |
7 | Amount | Negative | -250.00 | Rejected |
8 | Amount | Above the limit | Limit plus 0.01 | Rejected or routed to approval, per requirement |
Per the SIX QR-bill FAQ, a QR reference requires a QR-IBAN, and “A QR-IID exclusively contains values in the range 30000–31999.” The QR reference has 26 numeric characters followed by a check digit, and only CHF and EUR are allowed (SIX QR-bill FAQ, 2026).
A format check doesn't prove the account exists: SIX says its IBAN check “does not confirm the actual validity of the IBAN, but only its formal structure.” Test “formally correct, account doesn't exist” at the interface to core banking (see API testing). Real customer IBANs don't belong in test data; see GDPR-compliant test data for synthetic alternatives.
What is the right expected result for invalid input?
The right expected result is a controlled rejection: a clear message, no data change, and no internal details. NIST SP 800-53 Rev. 5, control SI-11, calls for error messages that give users the information they need to correct the problem without revealing information an attacker could exploit (NIST SP 800-53 Rev. 5, 2020).
Check these points every time:
The message names the field and the rule, such as “The IBAN must have 21 characters.”
The response contains no stack trace, no SQL fragment, and no server name.
No payment was stored and no balance changed.
The HTTP status code fits the cause: 400 for a malformed request, never 500.
The user can correct the input without re-entering every field.
A crash with an error page is a defect, even when the payment never went through. The test passes only when the rejection matches what you defined in step 5.
Where does negative testing end and security testing begin?
The line runs along the goal: enforcing business rules with selected invalid values on one side, finding exploitable weaknesses with attack patterns or mass-generated input on the other. MITRE lists improper input checking as CWE-20, ranked 18th in the 2025 CWE Top 25 (MITRE CWE Top 25, 2025).
Criterion | Negative testing | Security testing and fuzzing |
|---|---|---|
Goal | Business rule is enforced | Weakness is found |
Input | Representatives of invalid partitions | Attack patterns, mass-generated data |
Oracle | Defined error message | Crash, data leak, unexpected behavior |
Owner | Test team and business analysts | Security team, penetration testers |
What Autemos handles
Autemos automates these functional checks from datasets. You store the cases as a dataset in CSV, XLSX, or JSON: one row per invalid partition, with columns for the test value and the expected message. Autemos loads the values at runtime as typed variables. One-time-use data is flagged so parallel runs don't reuse it. Details: Autemos test data handling.
Autemos does not run penetration tests or fuzzing. SQL injection, cross-site scripting, or protocol fuzzing need a dedicated security tool or specialist provider.
Frequently asked questions
Is there an ISTQB definition of positive testing?
No, the official ISTQB glossary has no entry for positive testing (checked September 27, 2026). It is industry shorthand for tests with valid input along the intended flow. The CTFL syllabus uses the phrase “positive test cases” and places them first.
What are examples of negative test cases?
Typical cases feed a system values it should reject: a 20-character Swiss IBAN, wrong IBAN check digits, a QR reference paired with a regular IBAN, the currency USD on a QR-bill, or an amount of 0.00 or -250.00. Each case holds exactly one invalid value.
How many invalid values should you test per input field?
An input field needs at least one test per invalid partition, plus the values just past each boundary. An amount field with a lower and an upper limit has at least two invalid partitions. The CTFL syllabus counts invalid partitions toward 100% equivalence partitioning coverage.
Can negative tests be automated?
Yes, negative tests suit automation from datasets: one test flow, many data rows. Each row holds one invalid partition and the expected message; Autemos reads such datasets from CSV, XLSX, or JSON and runs them across web, mobile, API, and desktop.
Does security testing cover the same ground?
No, negative testing checks business input rules with selected invalid values, and security testing looks for exploitable weaknesses. They overlap at input handling (CWE-20). Penetration testing and fuzzing need their own tools and expertise.
Conclusion
Negative testing checks whether a system rejects invalid input in a controlled way. ISTQB defines it as use in a way that is not intended; positive testing has no glossary entry. 92% of the catastrophic failures Yuan et al. studied came from error handling, and the FCA case against Citigroup shows where a missing hard block let US$1.4bn of equities be sold by mistake.
Derive the cases from invalid partitions, boundaries, business rules, and likely errors. Test exactly one invalid partition per case and define the expected message before you run it. For Swiss payments: IBAN length, check digits, QR-IBAN, CHF or EUR, and amount limits, one at a time. If you want to automate datasets like these in Autemos, talk to our team.


