Multimodal Data
NCA-GENM · 45 questions
- A maple-sugar shack wants NVIDIA Riva ASR to recognize local terms such as “spile,” “niter,” and “hydrometer.” Staff prepare a word list plus example utterances. How should that material be treated for Riva ASR customization?
- A locksmith bench records spoken key-cut requests and types what was said for each clip. What is the correct unit of ASR customization data?
- An accordion-repair loft dumps a week of bench recordings with no typed lines and labels the dump “the Riva set.” What should happen before treating it as ASR customization data?
- A weathervane loft wants NVIDIA Riva TTS to read care notices aloud. A volunteer points at a pile of vane photographs as “the TTS data.” What should the team request instead?
- A ship-caulker records a long take that includes the phrases “oakum, then pitch.” One annotator marks when each phrase starts; another writes a single mood note on the whole file. Which annotation helps the speech recognizer learn from the clip?
- A parchment studio must choose whether to record its Riva ASR set over the scraping bench or in a quiet side room. The finished kiosk will listen on the busy floor. How should the recording-condition choice be framed?
- An indigo-vat floor has one talkative dyer and six quieter helpers, but the first ASR draft corpus contains only the dyer. What should the team flag before customizing Riva?
- A clay-pipe kiln writes how specialty terms such as “sagger” and “waster” should be pronounced for NVIDIA Riva ASR. How should that pronunciation lexicon be classified?
- A paper-marbling shop may later add a talking face using ACE with Riva and Audio2Face from official additional materials. What is the smart Domain 3 move for today’s audio–transcript pairs?
- A sign-painter kiosk must store what was heard, what was meant, and what should be spoken back for each customer turn. How should that record be treated as conversational pipeline data?
- A dry-stone yard has ASR transcripts and intent tags but no spoken-back lines for the kiosk. What does that mean for the conversational pipeline corpus?
- A millstone loft wants NLP to route “dress the stone” versus “buy flour,” but someone starts tagging hopper photographs instead. Where should the intent annotation go?
- A farrier stall’s ASR set is English, the intent table mixes languages, and the TTS scripts are written in yet another language. What should the candidate conclude about the three corpora?
- A gilder’s loft has leaf-and-frame photographs and no spoken turns. A contractor calls the album “the multimodal Riva set.” How should that claim be handled?
- A pewter bench records overlapping customers, and without turn boundaries the intent tagger cannot see where one request ends. What role do timestamps or turn segmentation play?
- A tinsmith row has a clean transcript, a blank intent field, and a spoken reply line left over from yesterday’s order. How should that example be diagnosed?
- A wainwright loft is only preserving axle talk for later reading and will not speak replies. How far should the speech data plan go?
- A malt-house floor snaps a floor-grain still and writes one English line that names what is in the frame. What is the CLIP data unit that describes?
- A sawmill office captions blade stills in a language the official CLIP skill path does not name for its English-prompt workflow. What should the team flag?
- A shingle-mill row shows a cedar shake, but the English line describes a slate roof from another bay. How should that CLIP row be treated?
- A lime-kiln camp has hundreds of pot stills and empty English caption fields. How should those rows be classified for CLIP?
- A coble yard has slogan lines such as “tarred bottom, oak strake” ready but no photographs. Can that slogan list be called a CLIP corpus?
- A last-maker can write what is actually on the bench or invent a poetic shop motto beside the still. Which English annotation belongs on a CLIP pair?
- A clog shop keeps stills in one drawer and English slogans in another with no shared identifier. What is required before those piles can become a CLIP paired corpus?
- A reed-organ loft can caption a still “walnut case, two manuals” or paste a three-page parts inventory. How should caption length and focus be treated?
- A rush-light shop intern scores caption–still rows with a spoken-word error count borrowed from the Riva booth. How should pairing quality for those rows be judged?
- A horn-button works has a still, a care sentence, and a tap-test clip that might belong to three different buttons. What must be true before anyone discusses early, late, or intermediate fusion on those files?
- A pin-mill row shows a brass pin in the still, but the audio clip is someone describing a steel coil from the next machine. What data problem should be marked?
- A needle-works row is planned as still + note + tap sound, but the microphone failed and no audio was captured. How should that row be treated in the multimodal corpus?
- A thimble shop loses the photo but still has the care sentence and a tiny click. Late fusion could run the remaining specialists; early concat cannot invent pixels. What should guide the next step?
- A fulling-mill intern concatenates a 4K still next to a three-word note and a 30-second clip with no shared clock or size layout. What should be flagged before calling the row early-fusion ready?
- A print-works desk wants a stills grader and a notes grader to vote at decision time, but the stills grader’s rows have no pictures. Can this yet be called a late-fusion dataset?
- A litho-stone loft plans to meet vision codes and text codes halfway for intermediate fusion. What data requirement comes first?
- A copperplate shop needs a label that is true only of the pair—this motto belongs on this plate—not of either file alone. Where does that annotation belong?
- A woodcut shop claims “late fusion of scent and stills” but never recorded any scent. What is the sound data response?
- An etching-bath intern wants to add a fourth unnamed product stream to every paired row “because multimodal means everything.” What scope should the corpus keep?
- An enamel kiln records a short firing clip and a typed question for an NVIDIA AI Blueprint with VIA. How should that material be treated as data?
- A wrought-iron shop sometimes has a spoken request and sometimes only a scroll still. What is the data side of modality orchestration for today’s example?
- A saddle-tree loft scripts greet → look up a tree → speak a reply, but the lookup table is empty. What should be named as the problem?
- A bridle loft planned still+speech rows, but the camera failed. How should today’s example be handled on the data side?
- A kelp-drying rack wants the system to “fetch the grade, then speak it,” but no grade table was built. Where should work stop?
- An eel-weir watch uses an NVIDIA AI Blueprint with VIA on a short clip plus a typed question. What are the data arms being orchestrated?
- An oyster-rake shed may later drive a face with ACE / Riva / Audio2Face. What data ask is appropriate at awareness depth?
- An ice-cutter pond office writes a polished greet / ask / show script but has no stills or captions. How should that script be filed?
- A packet-boat office argues about the next scripted agent action before anyone records the passenger still or the spoken request. What should happen first?