Half My Survey Responses Look Like Junk: How to Screen Them Without Wrecking Your Sample (2026)

·

Half My Survey Responses Look Like Junk: How to Screen Them Without Wrecking Your Sample (2026)

You chased responses for six weeks. You finally have 240. And when you open the file, something is wrong: a run of submissions that arrived within four minutes of each other, dozens of people who answered a twelve-minute questionnaire in ninety seconds, whole columns where somebody has ticked “4” forty times in a row, and an open-text box containing a sentence that answers a question you never asked.

This is the panic nobody warns you about, because it arrives after the stage everyone warns you about. The recruitment problem is over. The problem now is that you cannot tell which of your 240 rows are people.

The good news is that this is a solved problem with a defensible procedure, and running that procedure properly turns what feels like a disaster into one of the strongest paragraphs in your methodology chapter. The bad news is that the procedure has to be applied in a particular order, and the order matters more than the individual checks.

Why this is happening to you specifically

Two things changed. Student surveys are now distributed almost entirely through open web links — social media, WhatsApp groups, subreddits, university mailing lists — and an open link is scrapeable. And where a survey offers any incentive at all, even a prize draw, it acquires an economic value that has nothing to do with your research question.

The result is a caseload of three distinct problems that students routinely treat as one:

  • Automated submissions. Scripted entries that complete the form without reading it. They cluster in time, share technical fingerprints, and produce open-text answers that are generic or lifted from elsewhere.
  • Fraudulent human responses. A real person completing your survey repeatedly, or misrepresenting their eligibility, to get the incentive. Harder to detect than bots, because the behaviour looks human.
  • Careless responding. A genuine, eligible participant who stopped paying attention halfway through. Not dishonest, but the data are still not measuring what you think.

These need different responses. Deleting all three under one heading called “invalid responses” is what makes a marker ask questions.

Rule one: decide the rules before you look at the results

This is the single point that determines whether your screening reads as method or as manipulation.

Write your exclusion criteria down, with thresholds, before you run a single analysis on the outcome variables. Date the note. It can go in your analysis plan, your supervision record, or an email to your supervisor — the medium matters less than the timestamp. It is the same discipline that governs whether you remove outliers, and for the same reason: a rule chosen before you see its consequences is evidence, and the identical rule chosen afterwards is not.

If you have already looked — most people have, because the mess is what prompted the question — you have not lost the argument. Write the criteria now, apply them mechanically, and report the sequence honestly, including that the problem was noticed during data screening. Transparency recovers most of what pre-specification would have given you. Concealment recovers none of it.

The screening sequence, in order

A questionnaire sheet with an entire column of identical ticks in a straight vertical line
Straightlining is the cheapest signal to detect and the easiest to defend as a criterion.

1. Eligibility. Before anything technical, remove responses that fail your stated inclusion criteria — wrong population, wrong country, outside the age range, answered “no” to your screening question. This is not a judgement call and it is not contentious.

2. Completion. Decide your minimum. A response that stops before the first outcome measure contributes nothing; a response missing two items of a thirty-item scale is usually salvageable. Where the line falls depends on your analysis, and the consequences of the choice are covered in the guide to handling missing data.

3. Duration. Your survey platform records how long each person took. The conventional speeding rule is to flag anything completed in less than about a third of the median completion time, on the reasoning that nobody can read the items that fast. Calculate the median from your own data rather than importing a threshold, and state the figure you used.

4. Response patterns. The long-string index counts the longest run of identical consecutive answers within a scale. A participant who selects “4” for all twenty items of a scale containing reverse-worded items has told you they were not reading. This check is powerful precisely because reverse-worded items make it almost impossible to pass by accident — one of several reasons to include them, as the guide to designing a Likert scale questionnaire sets out.

5. Attention and consistency checks. If you built in an instructed item (“select ‘Strongly agree’ for this statement”), apply it now. If you did not, you can still use consistency: two items asking the same thing in opposite directions should not both be endorsed. A single failure is weak evidence and two are strong, which is why the standard advice is to include at least two and to pre-specify how many failures trigger exclusion.

6. Technical duplicates. Repeated submissions from the same address, implausible geolocation given your recruitment, and clusters of submissions seconds apart all point to the same source. Keep the first complete response and exclude the rest, and say which one you kept. Note that IP addresses are personal data under UK GDPR — the guide to survey tools and data protection covers what your ethics approval needs to say about collecting them.

7. Open-text plausibility. The strongest single signal, and the one no automated filter catches. A generic paragraph that answers nothing specific, or an answer irrelevant to the question, is the clearest evidence you will find. One free-text question exists to do this job even if you never analyse the answers.

How much can you remove before it becomes a problem?

Two separated stacks of questionnaire responses on a desk beside a magnifying glass and a screening checklist
The excluded pile is not waste. Described properly, it is evidence of rigour.

There is no percentage that is automatically too high. What matters is whether the exclusions are systematic and whether enough remains.

Check two things. First, that your remaining sample still meets the size your power analysis called for — and if it does not, say so and treat it as a limitation rather than hoping nobody recalculates. Second, that the excluded cases do not differ systematically from the retained ones on your demographics. If every excluded response came from one recruitment channel, that is worth a sentence, because it changes who your sample represents. The sampling methods guide explains why that shift matters more than the raw number lost.

And if screening leaves you genuinely short, reopening collection with a closed link, a screening question and an attention check is usually faster than defending an unusable dataset. Choosing the platform on its fraud-prevention features rather than its interface is worth a few minutes at that point — the comparison of academic survey tools covers which ones offer what.

What to write in the methodology chapter

Your screening deserves a short subsection of its own, and it should read as a sequence with numbers attached. Six sentences is enough.

State the criteria and when you set them. Give the thresholds you actually used, not the ones you read about. Report the number removed at each step, so the arithmetic runs from your raw total to your analysed total without a gap. Note whether excluded cases differed from retained ones. State the final N and, if you can, the direction of any effect the exclusions had.

Of 240 submissions, 18 failed the eligibility screen, 11 were incomplete before the first outcome measure, 24 completed in under a third of the median duration, 9 showed a long-string run exceeding 15 consecutive identical responses across reverse-worded items, and 6 were duplicate submissions from a single address. Criteria were set on 14 March, before any outcome analysis. The final analysed sample was 172. Excluded and retained cases did not differ significantly by age or gender.

That paragraph does more for your methodology mark than a clean dataset would have done, because it demonstrates something a clean dataset never gets the chance to.

Frequently asked questions

Should I say in my dissertation that I had bot responses?

Yes. It is a known and documented hazard of open-link recruitment in 2026, and describing how you detected and handled it demonstrates competence. Concealing it and reporting only the final N leaves an unexplained gap between your recruitment account and your sample size.

What if I did not include an attention check?

You can still screen on duration, long-string patterns, duplicates, internal consistency and open-text plausibility — five independent signals. Say plainly that no instructed-response item was included and treat it as a design limitation.

Can I exclude someone just for answering too fast?

On a pre-specified threshold applied to everyone, yes, and it is standard practice. Applied selectively after seeing whose answers you dislike, no. The defence is the rule, not the individual case.

Do I need ethics approval to exclude responses?

No, but your approved protocol should be consistent with what you did. If it specified an analysis sample and you have changed how it is defined, note it in your final report. Collecting IP addresses or geolocation to detect duplicates is the part that genuinely engages your data protection commitments.

Is careless responding the same as an outlier?

No, and conflating them is a common error. Careless responding is a question about whether a participant belongs in the dataset at all; an outlier is an unusual value from a participant who does belong. They are screened at different stages and justified on different grounds.

Should I remove people who failed only one attention check?

Usually not. One failure can be a misread. Pre-specify a threshold of two or more failures out of the checks you included, and report both the threshold and the number excluded by it.

My prize draw attracted the fraud. Should I have skipped the incentive?

Not necessarily — incentives raise response rates, which is its own problem to solve. The lesson is to pair an incentive with a closed or single-use link and to separate the prize-draw entry from the survey itself, so completing the questionnaire repeatedly gains nothing.

Can I just use secondary data next time?

For many questions, yes, and it removes this entire class of problem along with the recruitment one. The trade-off is that you answer the question the data were collected for. See desk-based versus primary data for how that decision usually goes.

The one thing to do now

Open a new document, write today’s date, and list your exclusion criteria with thresholds before you run anything else. Everything above depends on that note existing, and it takes four minutes.

Then, when you write the screening subsection, Tesify can turn your criteria and your counts into the methodology passage in the order examiners expect — the sequence, the thresholds, the attrition arithmetic and the comparison of excluded against retained cases — so the six weeks you spent collecting the data are represented by a paragraph that defends them.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *