What is document 'deduplication' in a Watson Discovery collection?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Todd Lammle: 'Imagine you're building a chatbot and this exact situation comes up — automatically identifying and filtering duplicate documents so they do not inflate search results is your go-to move. Deduplication in Watson Discovery detects and removes or filters duplicate document ingestions so the same content is not indexed multiple times; which would inflate search result counts and reduce quality. This is a classic Domain 3: Connect AI and External Knowledge concept you'll want locked in before exam day.'
Full explanation below image
Full Explanation
Deduplication in Watson Discovery detects and removes or filters duplicate document ingestions so the same content is not indexed multiple times; which would inflate search result counts and reduce quality. It is not related to intent examples; session variables; or repeated messages. The correct answer, "Automatically identifying and filtering duplicate documents so they do not inflate search results", directly satisfies the scenario because it aligns with watsonx Assistant's design principles and the specific capability being tested. The incorrect options ("Removing duplicate intent training examples", "Removing repeated session variables at conversation start", "Filtering repeated user messages within a session") may appear relevant but each misses a key requirement or introduces a step that is either unnecessary or belongs to a different workflow. Mastering the distinction between these approaches is essential for effective watsonx Assistant implementations and is a core focus of the Domain 3: Connect AI and External Knowledge section of the certification exam.