Question 1
What are best practices for ensuring reproducibility and versioning in a data science project?
Correct Answer:
Version control for code; data versioning; environment management; dependency pinning; logging; and documentation.
Explanation:
Reproducibility in data science hinges on capturing code, data, and the runtime environment, and making them versionable and well documented. Version control for code keeps a traceable history of changes, so you can understand how results evolved and revert to a previous state if needed. Data versioning ensures the exact dataset used (including preprocessing steps and transformations) can be retrieved later, preventing shifts in results when data changes. Environment management guarantees the same runtime setup—Python or R version, OS specifics, and a consistent set of libraries—so runs aren’t affected by subtle system differences. Dependency pinning locks specific library versions, protecting against drift that can alter behavior or results across runs. Logging captures the details of each experiment—parameters, seeds, data splits, and metrics—so you can reproduce a particular run precisely and compare results across attempts. Documentation ties everything together, explaining how to reproduce steps, what each parameter means, and how data flows through the pipeline. Together, these practices create a stable, auditable trail from code to results, which is essential for reliable replication, collaboration, and long-term maintenance.
Question 2
Which statement about a unique key is true?
Correct Answer:
A unique key ensures all values in a column are distinct and allows one NULL value.
Explanation:
A unique key enforces that the values in its column (or set of columns) are distinct across all rows. This means a given non-NULL value can appear at most once, so you can identify a row by its unique value. It also allows NULL values; in practice, many databases permit multiple NULLs because NULL represents unknown and NULLs are not considered equal to each other. So the core idea is that a unique key guarantees uniqueness of values and permits NULLs, unlike a primary key which cannot be NULL.
Question 3
What does TCL manage?
Correct Answer:
Controlling transactions, such as COMMIT and ROLLBACK.
Explanation:
Transaction control language focuses on managing groups of database operations as a single unit. It lets you start a transaction, perform multiple changes, and then decide whether to apply all of them or undo them. The main commands COMMIT and ROLLBACK are used to finalize changes or revert them if something goes wrong. This control is what preserves atomicity and consistency across operations, ensuring that either everything in the transaction takes effect or none of it does. Some systems also support SAVEPOINT to mark a point to which you can roll back within a transaction. In contrast, defining database schemas, querying data, or setting permissions belong to other SQL sublanguages, while TCL specifically handles transaction flow.
Question 4
Which statement best describes the difference between UNION and UNION ALL?
Correct Answer:
UNION removes duplicates and UNION ALL keeps duplicates, which is faster.
Explanation:
When combining results from multiple queries, the key idea is how duplicates are treated. UNION returns only distinct rows, removing any duplicates that appear across the combined results. UNION ALL keeps every row from both sides, including duplicates, so the same row can appear more than once if it comes from both queries. This difference in handling duplicates is what drives the typical performance gap: removing duplicates requires extra work to identify and drop duplicates, so UNION usually has more overhead and can be slower than UNION ALL. Use UNION when you need a unique set of rows; use UNION ALL when you want to preserve duplicates or when you want the faster, simpler operation and duplicates don’t matter for your result. For example, if one query returns (1, 'x') and the other also returns (1, 'x'), UNION would yield a single (1, 'x'), while UNION ALL would yield two copies of (1, 'x').
Question 5
Which statement correctly contrasts embedding and referencing in MongoDB?
Correct Answer:
Embedding stores related data as a nested document; Referencing stores related data in a separate collection
Explanation:
Embedding and referencing describe how related data is modeled in MongoDB. The correct statement is that embedding stores related data as a nested document, while referencing stores related data in a separate collection. Embedding puts the related data inside the parent document, which makes reads fast and atomic for the whole object and is great for tightly linked, one-to-few relationships. But it can blow up document size and create duplication if the related data is large or needs to be updated independently. Referencing keeps related data in separate documents across collections and uses identifiers to link them, which avoids duplication and supports more flexible, many-to-many relationships, but reads require extra queries or a lookup stage to assemble the full picture. The other statements misstate where the data lives: embedding is not in a separate collection, and referencing is not a nested document or stored in the same document.
Question 1
Exam overview

About this Exam

Prepare with the DDR Data Science Interview Practice Test practice quiz. This question bank includes 10 questions covering union, mongodb, many, data, and science. Use it to review important concepts, identify knowledge gaps, and build confidence for the related exam, course, or assessment.

More details

Additional Information

DDR Data Science Interview Practice Test

This practice set contains 10 questions from the matching question bank and focuses on union, mongodb, many, data, and science. Work through each question carefully, review the provided solutions, and revisit topics that need more study before your next attempt.

This is an independent study resource intended for practice and review; it is not an official examination or an endorsement by any organization named in the title.

Quiz information

Frequently Asked Questions

The complete question count is available after full access is unlocked.
No fixed duration is currently configured for this quiz.
Question explanations are included where they are available in the quiz content, helping you review the reasoning after answering.
Yes. You can retake the practice test again as you continue studying during your available access period.
After your access is confirmed, you can continue into the complete practice exam from this quiz flow.
Unless explicitly stated otherwise, this page provides independent practice material for study and exam preparation and is not the official examination itself.
Keep studying

Related Questions