An oyster-rake shed may later drive a face with ACE / Riva / Audio2Face. What data ask is appropriate at awareness depth?
Select an answer to reveal the explanation.
Short Explanation
If a face will be driven from speech, you need paired speech and face recordings. Ask for those takes; don’t build the whole microservice map in this domain.
Full Explanation
Avatar speech-in and face-out paths consume paired speech and target-face recordings when that path is in scope. At multimodal-data awareness depth, the ask is for those paired takes—not microservice assembly. CLIP still–caption rows alone do not supply face-drive pairs.