ETExamTower
Q20Applications of Foundation Models

A company is introducing a mobile app that helps users learn foreign languages. The app makes text more coherent by calling a large language model (LLM). The company collected a diverse dataset of text and supplemented the dataset with examples of more readable versions. The company wants the LLM output to resemble the provided examples. Which metric should the company use to assess whether the LLM meets these requirements?

← → navigate · a answer
Community votes
C
100% (6)
A
0% (0)
B
0% (0)
D
0% (0)
Discussion · 5
C 6
The ROUGE (Recall-Oriented Understudy for Gisting Evaluation) score is widely used to measure the similarity between generated text and a set of reference texts. Since the company wants the LLM's output to resemble the provided readable examples, ROUGE is the most appropriate metric. ROUGE compares the LLM-generated text with the human-provided reference texts by evaluating n-gram overlap, precision, recall, and F1 score, making it a great choice for text coherence and readability assessment.
C 1
Since the company wants the LLM output to resemble the provided examples in terms of coherence and readability, ROUGE score is the best metric for this evaluation.
C 1
The correct answer is C. ROUGE score measures how well generated text matches reference examples.
C 1
he most suitable metric to assess whether the LLM output resembles the provided examples of more readable text is: C. Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score The ROUGE score is commonly used for evaluating the quality of text summarization and machine-generated text by comparing it to a set of reference texts. It measures how well the generated text matches the provided examples in terms of content and coherence. Specifically, ROUGE scores focus on the overlap of n-grams, word sequences, and word pairs between the generated text and the reference texts, making it ideal for this use case.
C 1
C is the correct answer