A transit analytics team caches several huge AVL history tables in Spark memory just in case, and executors begin failing with out-of-memory errors. What lesson should the program lead take?
Select an answer to reveal the explanation.
Short Explanation
Think of Spark memory like a city garage with a fixed number of stalls. Stuff every spare plow and bus in just in case, and the next vehicle has nowhere to go. Caching helps only when the data fits the space you actually have.
Full Explanation
Spark's in-memory cache accelerates repeated reads, but executors have finite heap and storage memory. Caching oversized or unnecessary datasets can exhaust that capacity and trigger out-of-memory failures. Civic teams should plan what to persist, estimate footprint, and cache selectively rather than assuming memory is unlimited.