Evaluation Framework · Four Stages
Question bank, execution, scoring and aggregation as one landscape evaluation framework.
6380 views
Start from this prompt
Question bank, execution, scoring and aggregation as one landscape evaluation framework.
Try Kimi DesignCreate a single-page landscape evaluation framework diagram at about 16:7, styled as a two-column academic-paper figure. Let it flow from Question Bank through Execution and Scoring to Aggregation. Question Bank contains question sources, type categories, difficulty tiers, and metadata for every question. Execution contains the evaluated model API, concurrent scheduling, retries, timeouts, and raw-output retention. Scoring has parallel rule-based and model-based paths feeding a consistency-check box; inconsistent samples go to human arbitration by dashed connector. Aggregation covers results by dimension, difficulty tier, and the final leaderboard. Archive raw outputs, linking the retention box to Scoring with a dashed arrow so scoring reruns reuse them. Use a white ground with several large, lightly tinted section panels. Give every section a lowercase letter index and name. Use rounded rectangles and a restrained set of contrasting hues: inputs and raw data share one hue family, intermediate outputs another, processing modules another, and actions or checks another; matching types match. Solid arrows carry data within sections, while dashed arrows transfer data between sections. Keep the page to the diagram, with dense, precise, content-rich labels and bold key-module names; use flat fills, plain surfaces, and generic unbranded styling.
Add confidence intervals
Shows a confidence interval next to each score on the final leaderboard, without changing the aggregation flow.
Try Kimi DesignCreate a single-page landscape evaluation framework diagram at about 16:7, styled as a two-column academic-paper figure. Let it flow from Question Bank through Execution and Scoring to Aggregation. Question Bank contains question sources, type categories, difficulty tiers, and metadata for every question. Execution contains the evaluated model API, concurrent scheduling, retries, timeouts, and raw-output retention. Scoring has parallel rule-based and model-based paths feeding a consistency-check box; inconsistent samples go to human arbitration by dashed connector. Aggregation covers results by dimension, difficulty tier, and the final leaderboard. Archive raw outputs, linking the retention box to Scoring with a dashed arrow so scoring reruns reuse them. Use a white ground with several large, lightly tinted section panels. Give every section a lowercase letter index and name. Use rounded rectangles and a restrained set of contrasting hues: inputs and raw data share one hue family, intermediate outputs another, processing modules another, and actions or checks another; matching types match. Solid arrows carry data within sections, while dashed arrows transfer data between sections. Keep the page to the diagram, with dense, precise, content-rich labels and bold key-module names; use flat fills, plain surfaces, and generic unbranded styling.
Switch to blue-gray
Retints the section panels from a neutral light tint to blue-gray, keeping nodes and connectors as-is.
Try Kimi DesignCreate a single-page landscape evaluation framework diagram at about 16:7, styled as a two-column academic-paper figure. Let it flow from Question Bank through Execution and Scoring to Aggregation. Question Bank contains question sources, type categories, difficulty tiers, and metadata for every question. Execution contains the evaluated model API, concurrent scheduling, retries, timeouts, and raw-output retention. Scoring has parallel rule-based and model-based paths feeding a consistency-check box; inconsistent samples go to human arbitration by dashed connector. Aggregation covers results by dimension, difficulty tier, and the final leaderboard. Archive raw outputs, linking the retention box to Scoring with a dashed arrow so scoring reruns reuse them. Use a white ground with several large, lightly tinted section panels. Give every section a lowercase letter index and name. Use rounded rectangles and a restrained set of contrasting hues: inputs and raw data share one hue family, intermediate outputs another, processing modules another, and actions or checks another; matching types match. Solid arrows carry data within sections, while dashed arrows transfer data between sections. Keep the page to the diagram, with dense, precise, content-rich labels and bold key-module names; use flat fills, plain surfaces, and generic unbranded styling.
Add section icons
Places a small icon beside the name of each of the four sections, next to the lowercase letter index.
Try Kimi DesignCreate a single-page landscape evaluation framework diagram at about 16:7, styled as a two-column academic-paper figure. Let it flow from Question Bank through Execution and Scoring to Aggregation. Question Bank contains question sources, type categories, difficulty tiers, and metadata for every question. Execution contains the evaluated model API, concurrent scheduling, retries, timeouts, and raw-output retention. Scoring has parallel rule-based and model-based paths feeding a consistency-check box; inconsistent samples go to human arbitration by dashed connector. Aggregation covers results by dimension, difficulty tier, and the final leaderboard. Archive raw outputs, linking the retention box to Scoring with a dashed arrow so scoring reruns reuse them. Use a white ground with several large, lightly tinted section panels. Give every section a lowercase letter index and name. Use rounded rectangles and a restrained set of contrasting hues: inputs and raw data share one hue family, intermediate outputs another, processing modules another, and actions or checks another; matching types match. Solid arrows carry data within sections, while dashed arrows transfer data between sections. Keep the page to the diagram, with dense, precise, content-rich labels and bold key-module names; use flat fills, plain surfaces, and generic unbranded styling.