DA0-002 Robots.txt Practice Question
Which THREE are best practices for acquiring data via web scraping? (Select exactly 3)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Respect robots.txt
Option B is correct because robots.txt is the standard mechanism (the Robots Exclusion Protocol) by which site owners declare which paths crawlers may access, and honoring it keeps scraping ethical and within the site's stated permissions. Option C is correct because setting a descriptive User-Agent header identifies your crawler and provides contact information, letting administrators recognize and reach you rather than blocking you as an anonymous bot. Option E is correct because limiting request rate (e.g., throttling to a modest number of requests per second and honoring Retry-After on 429/503 responses) avoids overloading servers and reduces the chance of being rate-limited or banned. Option A does not belong because rotating multiple IP addresses to evade blocks is an evasion technique, not a best practice, and can violate terms of service. Option D does not belong because ignoring terms of service and scraping indiscriminately is unethical and often unlawful, contrary to responsible scraping.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use multiple IP addresses
Why it's wrong here
Rotating source IP addresses evades rate limiting and blocking, which breaches most sites' terms and risks legal or access consequences. It is tempting because distributed requests can raise throughput and avoid throttling during large crawls. Respecting robots.txt, throttling request rates, and honouring terms of service are the accepted practices instead.
- ✓
Respect robots.txt
Why this is correct
Honouring robots.txt satisfies the legal and ethical constraint of scraping only permitted paths. The file declares which directories crawlers may access, so parsing it before requesting pages prevents the crawler from breaching site owner directives and reduces the risk of being blocked or facing legal action.
- ✓
Identify yourself with a user-agent
Why this is correct
Supplying a descriptive user-agent satisfies the traceability constraint, letting site owners identify the crawler and contact its operator. This distinguishes legitimate automated acquisition from anonymous abusive traffic, so administrators can whitelist or throttle the client rather than blocking it outright.
- ✗
Scrape all data without regard to terms
Why it's wrong here
Scraping without regard to terms violates site policies, robots.txt directives, and potentially copyright or computer-misuse law. It is tempting because unrestricted collection maximises dataset coverage quickly. Best practice instead requires reviewing and honouring terms of service, robots.txt, and applicable data-protection obligations before acquisition.
- ✓
Limit request rate
Why this is correct
Throttling request rate satisfies the constraint of not degrading the target site's availability. Pacing requests below the server's capacity prevents denial-of-service effects and avoids triggering rate-limit or IP-block responses, keeping the scraping pipeline stable and considerate of shared infrastructure.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.