Courseiva

NCA-GENL Data Analysis and Visualization Practice Question

A data scientist is analyzing 2 million document embeddings from a RAG corpus on a single NVIDIA GPU. A full pairwise cosine similarity matrix would require roughly 16 TB of memory, which is infeasible. They need to identify near-duplicate documents and visualize cluster density without materializing the full matrix. Which approach is most appropriate?

⚠ Common exam trap

The trap here is treating dimensionality reduction as a substitute for similarity search, when projection distorts the very distances you are trying to measure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use FAISS with an IVF-PQ index to perform approximate nearest-neighbor search, then plot a 2D projection colored by local density.

At 2 million vectors, the full similarity matrix is memory-infeasible, so an approximate method is required. FAISS IVF-PQ combines inverted-file partitioning with product quantization to search a compressed index on a single GPU, returning near-duplicates efficiently. Plotting a 2D projection colored by local neighbor density then exposes cluster structure and duplicate hotspots. Exact full-matrix computation, PCA-to-2D distances, and norm sorting all fail on memory or accuracy grounds.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Sort embeddings by their L2 norm and treat documents with similar norms as duplicates.

    Why it's wrong here

    The L2 norm of an embedding is a scalar magnitude and carries no directional information about semantic content. Two unrelated documents can share a norm while two near-duplicates can differ slightly in magnitude. Using norm similarity as a duplicate proxy is technically incorrect and will generate many false positives.

  • ✗

    Compute the full 2M x 2M cosine similarity matrix in float32 on the GPU to preserve exact distances.

    Why it's wrong here

    Storing a 2M x 2M float32 matrix requires about 16 TB, far beyond any single GPU's memory. Even tiling the computation does not avoid the storage problem if the matrix is materialized. This approach is infeasible at the stated scale and ignores the requirement to avoid the full matrix.

  • ✓

    Use FAISS with an IVF-PQ index to perform approximate nearest-neighbor search, then plot a 2D projection colored by local density.

    Why this is correct

    FAISS IVF-PQ compresses vectors into product-quantized codes and searches only a subset of inverted lists, so it finds near-duplicates without ever building the full similarity matrix. Coloring a 2D projection by local neighbor density reveals cluster structure and duplicate hotspots. This scales to millions of vectors on one GPU and directly answers both the duplicate and density questions.

  • ✗

    Reduce dimensionality to 2D with PCA first, then compute exact pairwise distances in the 2D space to find duplicates.

    Why it's wrong here

    PCA to two dimensions discards most variance, so documents that are near-duplicates in the original embedding space can collapse or separate unpredictably. Exact distances in 2D are not faithful to the high-dimensional geometry, producing false duplicates and missed ones. This sacrifices accuracy for speed in a way that undermines the analysis.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.