Courseiva

DA0-002 Robots.txt Practice Question

Which THREE are best practices for acquiring data via web scraping? (Select exactly 3)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Respect robots.txt

Option B is correct because robots.txt is the standard mechanism (the Robots Exclusion Protocol) by which site owners declare which paths crawlers may access, and honoring it keeps scraping ethical and within the site's stated permissions. Option C is correct because setting a descriptive User-Agent header identifies your crawler and provides contact information, letting administrators recognize and reach you rather than blocking you as an anonymous bot. Option E is correct because limiting request rate (e.g., throttling to a modest number of requests per second and honoring Retry-After on 429/503 responses) avoids overloading servers and reduces the chance of being rate-limited or banned. Option A does not belong because rotating multiple IP addresses to evade blocks is an evasion technique, not a best practice, and can violate terms of service. Option D does not belong because ignoring terms of service and scraping indiscriminately is unethical and often unlawful, contrary to responsible scraping.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use multiple IP addresses

    Why it's wrong here

    Rotating source IP addresses evades rate limiting and blocking, which breaches most sites' terms and risks legal or access consequences. It is tempting because distributed requests can raise throughput and avoid throttling during large crawls. Respecting robots.txt, throttling request rates, and honouring terms of service are the accepted practices instead.

  • ✓

    Respect robots.txt

    Why this is correct

    Honouring robots.txt satisfies the legal and ethical constraint of scraping only permitted paths. The file declares which directories crawlers may access, so parsing it before requesting pages prevents the crawler from breaching site owner directives and reduces the risk of being blocked or facing legal action.

  • ✓

    Identify yourself with a user-agent

    Why this is correct

    Supplying a descriptive user-agent satisfies the traceability constraint, letting site owners identify the crawler and contact its operator. This distinguishes legitimate automated acquisition from anonymous abusive traffic, so administrators can whitelist or throttle the client rather than blocking it outright.

  • ✗

    Scrape all data without regard to terms

    Why it's wrong here

    Scraping without regard to terms violates site policies, robots.txt directives, and potentially copyright or computer-misuse law. It is tempting because unrestricted collection maximises dataset coverage quickly. Best practice instead requires reviewing and honouring terms of service, robots.txt, and applicable data-protection obligations before acquisition.

  • ✓

    Limit request rate

    Why this is correct

    Throttling request rate satisfies the constraint of not degrading the target site's availability. Pacing requests below the server's capacity prevents denial-of-service effects and avoids triggering rate-limit or IP-block responses, keeping the scraping pipeline stable and considerate of shared infrastructure.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.