Evaluate and Improve Model Responses
Build a source-backed workbook to compare, score, and improve AI responses.
Introduction
30 Second Summary
Polished answers can feel trustworthy at first glance. A careful review reveals whether that confidence is deserved.
In this project, you will build an AI Trainer Evaluation Workbook that turns a quick preference into a repeatable judgment backed by evidence. You'll use pairwise comparison in Google Sheets to score anonymized ChatGPT and Claude responses against clear success criteria.
What You'll Build
Picture sharing a view-only workbook where a reviewer can follow one model comparison from the original prompt to a sourced final verdict.
By the end of this project, you'll have:
- A side-by-side comparison that lets you judge two anonymized model responses without product names influencing the decision.
- A five-criterion rubric with clear 1-to-3 anchors that makes every score auditable.
- A view-only work sample that lets a reviewer trace the reasoning behind your final decision.
- Secret Mission: Add a missing-evidence test case that returns Insufficient Evidence when the required source passage is absent.
Are there any prerequisites?
No previous AI evaluation experience is needed. The guide includes setup for every tool plus a ChatGPT-only path if Claude is unavailable.
Before We Start
This is your moment to commit to the work sample before any hands-on setup begins. A clear goal keeps every later judgment tied to the AI Trainer role you want.
Set Up Your Evaluation Workspace
Careful AI evaluation depends on a traceable record. Scattered browser tabs make that record easy to lose.
This step creates one workspace in Google Chrome. It combines a Google Sheets workbook with two AI chat surfaces.
You will test ChatGPT plus Claude with one simple prompt. That gives you two visible replies before the real comparison begins.
In this step, get ready to:
- Install or open Google Chrome.
- Create the three-tab evaluation workbook.
- Verify two AI chat workspaces with the same prompt.
Install and open Google Chrome
Google Chrome keeps the workbook plus both chats within one browser workspace. Start by checking whether it is already installed.
- Press Cmd+Space on macOS or the Windows key on Windows.
- Type Google Chrome into the system search bar.
Google Chrome appears in the search results when it is already installed. Choose the path that matches what you see.
✔️ Google Chrome appears
- Press Enter to open Google Chrome.
Great, Google Chrome is open. Your browser workspace is ready.
ⓧ Google Chrome is missing
- Use an available browser to visit the official Google Chrome page.
- Use your device's official app store if the product page does not offer an installer.
- Follow the installation prompts for your device.
Google Chrome is installed once the installation flow finishes.
- Return to system search with Cmd+Space on macOS or the Windows key on Windows.
- Type Google Chrome into the search bar.
Google Chrome should now appear in the search results.
- Press Enter to open Google Chrome.
Google Chrome is now open. You have the browser needed for the rest of the project.
Still unable to open Chrome?
Restart your device if the installed browser does not appear in system search.
Get help with the installation issue by sharing your device type plus what happens when you search for Chrome. Help me troubleshoot my Google Chrome installation.
Create the evaluation workbook
The workbook keeps every later prompt plus judgment in one reviewable place. You will begin with three blank tabs so each stage has a clear home.
- Create a new Chrome tab by clicking the plus-shaped control in the tab bar.
You should see a blank browser tab with the address bar ready for input.
- Enter https://sheets.google.com/create in the address bar.
- Press Enter.
Google Sheets opens a blank spreadsheet or asks you to access a Google Account.
- Create a Google Account from the account screen if you do not already have one.
- Sign in to the Google Account you want to use for this project.
The blank spreadsheet opens after account access is complete. This is the file that becomes your evaluation workbook.
- Click the workbook title in the top-left corner.
- Type AI Trainer Evaluation Workbook.
The new workbook name is now visible in the title field.
- Press Enter to save the workbook name.
You should see AI Trainer Evaluation Workbook at the top of the spreadsheet.
- Double-click the first sheet tab at the bottom of the workbook.
- Type Rubric.
The first tab name is ready to save.
- Press Enter to save the tab name.
You should see a blank tab named Rubric.
- Click the plus-shaped control beside the sheet tabs.
- Double-click the new sheet tab.
The new blank tab name is now editable.
- Type Evaluations.
- Press Enter to save the tab name.
You should now see blank tabs named Rubric plus Evaluations.
- Click the plus-shaped control beside the sheet tabs.
- Double-click the new sheet tab.
The third blank tab name is now editable.
- Type Portfolio Summary.
- Press Enter to save the tab name.
Your three blank tabs are ready. The workbook now separates rubric design from evaluation records plus the final reviewer summary.
Workbook tabs not matching?
- Press Enter after typing each tab name.
- Keep exactly three blank tabs named Rubric, Evaluations, and Portfolio Summary.
Get help by describing the tab names currently visible at the bottom of your workbook. Help me fix my Google Sheets tab setup.
Verify two AI chat workspaces
A short verification prompt checks that each chat accepts instructions plus returns a visible answer. You will use the same prompt in both places so the setup check stays consistent.
These sign-ins only establish access for this project. Keep confidential employer material plus personal information out of both chats.
- Create a new Chrome tab by clicking the plus-shaped control in the tab bar.
You should see a blank browser tab with the address bar ready for input.
- Enter https://chatgpt.com in the address bar.
- Press Enter.
ChatGPT opens in the browser. It may ask you to create an account or sign in before saving the chat.
- Complete the account creation or sign-in prompts shown by ChatGPT.
You should see a chat message box once account access is complete.
- Copy the exact verification prompt below:
Reply with READY and one sentence defining an AI evaluation.
Why use a verification prompt?
The required word READY makes successful instruction following easy to spot. The sentence confirms that the chat can return a complete response.
- Paste the verification prompt into the ChatGPT message box.
- Send the message.
You should see a visible reply containing READY plus one sentence that defines an AI evaluation.
No ChatGPT reply?
- Confirm that ChatGPT shows you as signed in.
- Send the verification prompt again if the first request remains unanswered.
Share what you see without including account details. Help me troubleshoot my ChatGPT verification chat.
- Create another Chrome tab by clicking the plus-shaped control in the tab bar.
You should see another blank browser tab with the address bar ready for input.
- Enter https://claude.ai in the address bar.
- Press Enter.
Claude requires users to be at least 18 years old. Access also depends on being in a supported location.
Choose the path that matches the page you see. The fallback preserves two separate candidate chats without blocking the project.
✔️ Claude is available
- Complete the account creation or sign-in prompts shown by Claude.
You should see a chat message box once account access is complete.
- Paste the same verification prompt from your clipboard into the Claude message box.
- Send the message.
You should see a second visible reply containing READY plus one sentence that defines an AI evaluation.
No Claude reply?
- Confirm that Claude shows you as signed in.
- Send the verification prompt again if the first request remains unanswered.
Share what you see without including account details. Help me troubleshoot my Claude verification chat.
ⓧ Claude is unavailable
Use a second ChatGPT conversation when Claude access is unavailable. Keeping the conversations separate still gives you two candidate workspaces.
- Return to the ChatGPT tab from earlier.
- Click New Chat.
You should see a separate empty ChatGPT conversation.
- Paste the same verification prompt from your clipboard into the message box.
- Send the message.
You should see a second visible reply containing READY plus one sentence that defines an AI evaluation.
Second chat not separate?
- Return to the ChatGPT conversation list.
- Select New Chat before resending the verification prompt.
Get help by describing the two conversations without sharing account details. Help me create two separate ChatGPT verification chats.
Before you make the final check, do you expect the workbook plus both chat replies to be ready at the same time?
- Return to the Google Sheets tab from earlier.
- Confirm the workbook title is AI Trainer Evaluation Workbook.
- Confirm the tabs named Rubric, Evaluations, and Portfolio Summary are blank.
- Return to the first ChatGPT verification chat.
- Confirm the verification prompt plus its visible reply are present.
- Return to the Claude verification chat or the second ChatGPT verification chat.
- Confirm the same verification prompt plus a second visible reply are present.
You should see the blank workbook open successfully. You should also see two visible AI replies to the same verification prompt.
Workspace checkpoint
- The Google Sheets workbook is named AI Trainer Evaluation Workbook with blank tabs named Rubric, Evaluations, and Portfolio Summary.
- The ChatGPT verification chat contains the prompt plus a visible reply.
- The Claude verification chat or second ChatGPT verification chat contains the same prompt plus a visible reply.
Your workspace now keeps the source material plus decisions together. Next up, you will use both chats to make an instinctive model comparison.
Compare Responses Instinctively
Your Google Sheets workbook is ready. Your ChatGPT conversation can now produce the first candidate response.
A pairwise comparison asks you to judge two answers to the same prompt. A quick preference can feel convincing. However, another evaluator may struggle to reproduce it from one informal reason.
In this step, get ready to:
- Generate two answers to the same vague prompt.
- Anonymize the replies as Candidate A and Candidate B.
- Record an instinctive winner with one first-impression reason.
Generate two responses to the same prompt
Using the same prompt keeps the request consistent. Any differences you notice come from how each model responds.
- Copy this shared prompt for both model conversations:
Explain why Earth has seasons in a clear and helpful way.
Why Use the Same Prompt?
Both models receive the same request. The broad wording gives each model room to choose its own explanation.
- Switch back to the ChatGPT conversation from the setup step.
- Paste the prompt into the message box.
- Send the message.
You will see one explanation of why Earth has seasons. Let the full response finish before moving on.
- Return to the Claude conversation from earlier. Use the second ChatGPT conversation from earlier if you are following the fallback.
- Paste the same prompt into the message box.
- Send the message.
Good progress. You now have two raw answers to the same question.
Missing a Complete Response?
- Resend the exact prompt in the affected conversation.
- Wait for the response to finish before copying it.
If a conversation still does not produce a usable reply, help me troubleshoot the missing model response.
Build the anonymized comparison
Anonymized labels keep product reputation out of your first judgment. The response text becomes the only thing you compare.
- Return to the first ChatGPT response.
- Copy the complete response.
- Switch back to AI Trainer Evaluation Workbook.
- Select the Evaluations tab at the bottom of the workbook.
- Enter Prompt in A1.
- Enter Explain why Earth has seasons in a clear and helpful way. in A2.
- Enter Candidate A in A4.
- Enter Candidate B in B4.
Keep the Candidates Anonymous
The neutral labels keep the comparison focused on the answers. Product names stay in the model conversations instead of appearing beside either candidate.
- Double-click A5 to edit the cell.
- Paste the first response unchanged into A5.
- Return to the Claude response. Use the second ChatGPT response if you are following the fallback.
- Copy the complete second response.
- Switch back to the Evaluations tab.
- Double-click B5 to edit the cell.
- Paste the second response unchanged into B5.
Your workbook now shows Candidate A beside Candidate B. Neither label reveals which product produced the response.
Record your first impression
An instinctive judgment captures your immediate preference before structured grading changes how you read the answers. This creates a baseline you can examine later.
Before you reread the pair, do you expect your preference to be obvious or close? Hold that prediction in mind.
- Read Candidate A once from top to bottom.
- Read Candidate B once from top to bottom.
- Enter Instinctive Winner in A8.
- Enter A, B, or Tie in B8.
- Enter First-Impression Reason in A10.
- Write one informal sentence in A11 explaining why you made that choice.
- Leave the Rubric tab blank.
- Leave the Portfolio Summary tab blank.
Before you inspect the finished comparison, do you think another evaluator could reproduce your choice from that one reason alone? Keep your answer in mind.
- Read the Evaluations tab from top to bottom.
You should see the vague prompt, Candidate A, Candidate B, your instinctive winner, and one informal reason. You should see no criterion-level scores or factual verification.
Why This Shortfall Matters
The workbook shows what you chose. It does not provide separate evidence for the qualities that shaped your choice.
Another evaluator can read your reason. They cannot repeat a scoring process that does not exist yet.
You have captured a genuine first-impression baseline. Next, you will make your evaluation process consistent enough for someone else to follow.
Build Your Scoring Rubric
Your instinctive comparison captured a genuine first impression. However, one informal reason gives another evaluator no consistent way to reproduce your choice.
A scoring rubric turns quality into separate dimensions with shared anchors. You will define those dimensions before collecting two fresh responses to a tighter prompt.
In this step, get ready to:
- Define five evaluation criteria with shared 1-to-3 scoring anchors.
- Generate two fresh responses to a constrained seasons prompt.
- Prepare an anonymized comparison with a blank scorecard.
Define the scoring anchors
The Rubric tab in Google Sheets will hold the same scale for every criterion. Shared anchors keep a score of 2 from meaning something different to each evaluator.
- Switch back to the AI Trainer Evaluation Workbook from earlier.
- Select the Rubric tab.
- Select cell A1.
- Fill the rubric grid by pasting this tab-separated content:
Criterion 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
Factual Accuracy 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
Instruction Following 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
Relevance 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
Clarity for the Audience 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
Source Quality 1 = Major Failure 2 = Partly Meets the Requirement 3 = Fully Meets the Requirement
How Does This Rubric Work?
Each row isolates one dimension of quality. This prevents a clear writing style from hiding a factual or instructional failure.
- Factual Accuracy checks whether the claims are correct.
- Instruction Following checks whether every prompt requirement is satisfied.
- Relevance checks whether the answer stays focused.
- Clarity for the Audience checks whether the intended reader can understand the answer.
- Source Quality checks whether the included link supports the answer.
The repeated anchors give every number a consistent meaning. That consistency makes the final judgment easier to audit.
- Review cells A1:D6.
You should see five criteria with the same three scoring levels. That gives another evaluator a shared measuring stick for the comparison.
Rubric Pasted Into One Cell?
- Undo the paste in the Rubric tab.
- Select cell A1 again.
- Copy the tab-separated content directly from the block above.
- Paste it into the selected cell.
Ask for help if the content still stays in one cell: Help me place tab-separated rubric content into separate Google Sheets cells.
Generate constrained responses
A tighter prompt exposes whether each response follows specific requirements. ChatGPT and Claude should receive the same wording so prompt differences cannot influence the comparison.
- Prepare the exact request by copying this prompt:
Explain why Earth has seasons to a 12-year-old in 100 words or fewer. Include Earth's axial tilt, explain that distance from the Sun is not the cause, and include one reliable source link.
Why Constrain the Prompt?
The audience requirement tests clarity. The word limit tests focus.
The required claims test factual coverage. The requested link makes source quality visible in the pairwise comparison.
- Switch back to the ChatGPT seasons conversation from earlier.
- Paste the constrained prompt into the message box.
- Send the prompt.
- Wait for the fresh response to finish.
- Switch back to the Claude seasons conversation or the second ChatGPT conversation from earlier.
- Paste the same constrained prompt into the message box.
- Send the prompt.
- Wait for the second fresh response to finish.
You should see two new answers that address the same audience, length limit, factual requirements, and sourcing request. You now have a fairer pair of responses for rubric-based grading.
Missing One Fresh Response?
Use the second ChatGPT conversation from earlier if Claude is unavailable. Keep the conversations separate so each response comes from an independent chat.
Get help with the fallback workflow: Help me generate two independent responses with separate ChatGPT conversations.
Prepare the blank scorecard
Anonymized labels keep your attention on the response content. A blank scorecard also preserves the boundary between rubric design and evidence-backed grading.
- Choose one of the two fresh responses as Candidate A.
- Copy the chosen response.
- Switch back to the AI Trainer Evaluation Workbook.
- Select the Evaluations tab.
- Select the first empty row below your informal reason.
- Enter Constrained Prompt in the first cell of that row.
- Paste the exact constrained prompt into the cell to the right.
You should see the new prompt below the instinctive comparison. The original prompt and first-impression decision remain unchanged above it.
- Enter Candidate A in the first cell of the next empty row.
- Paste the chosen response into the cell to the right.
You should see the full fresh response beside Candidate A.
Keep the Candidates Anonymous
Use only Candidate A and Candidate B in the workbook. Leave the product names in the browser conversations.
Anonymous labels reduce the chance that product preference influences your scores.
- Return to the other model's fresh response.
- Copy the response.
- Switch back to the Evaluations tab.
- Enter Candidate B in the first empty row below Candidate A.
- Paste the copied response into the cell to the right.
You should now see two fresh responses with anonymous labels. No product name should appear beside either candidate.
- Select an empty area beside the two fresh responses.
- Add the blank five-criterion scorecard by pasting this tab-separated content:
Criterion Candidate A Score (1-3) Candidate B Score (1-3)
Factual Accuracy
Instruction Following
Relevance
Clarity for the Audience
Source Quality
Why Leave the Scores Blank?
The scorecard records the same five dimensions defined in the Rubric tab. The empty fields keep this step focused on building the evaluation structure.
Evidence checking comes before grading. This prevents confident wording from earning points before its claims are verified.
- Leave every candidate score cell empty.
Before you inspect the finished setup, do you expect the new scorecard to contain any numbers?
- Read the Evaluations tab from top to bottom.
- Confirm that the vague prompt comparison remains at the top.
- Check that the constrained prompt appears below the informal reason.
- Check that both fresh candidate responses are present.
- Check that all five candidate score fields remain blank.
- Switch back to the Rubric tab.
- Confirm that all five criteria have the three shared anchors.
You should see a complete five-criterion rubric plus a fresh anonymized comparison that is ready for grading. The blank scores are intentional because no factual verification has happened yet.
Scorecard Missing a Criterion?
Compare the scorecard with the Rubric tab. Both places should list the same five criteria in the same order.
Ask for help if a row or candidate column is missing: Help me check my five-criterion evaluation scorecard in Google Sheets.
Your rubric and fresh comparison are ready. Next, you will verify the factual claims before choosing scores, confidence, and a winner.
Grade with NASA Evidence
Your constrained prompt now sits beside two anonymized responses in the Google Sheets workbook. The five-criterion rubric gives you a shared scoring language.
Fluent wording can still earn an instinctive preference without proving the central claims. NASA evidence lets you separate factual accuracy from confidence.
In this step, get ready to:
- Verify the central seasons claims against NASA evidence.
- Score both candidates using the five existing criteria.
- Write an evidence-backed adjudication note.
Verify the central claims with NASA
An authoritative source gives every evaluator the same factual baseline. You will use NASA Space Place to check the two claims at the center of the prompt.
- Switch back to Google Chrome.
- Open a new browser tab.
- Load NASA Space Place by entering https://spaceplace.nasa.gov/seasons/ in the address bar.
- Confirm that the page heading reads What Causes the Seasons?.
- Read the explanation of Earth's tilted axis.
You'll see that Earth's tilt changes how directly sunlight reaches each hemisphere during the year. You'll also see that changing Earth-Sun distance does not cause the seasons.
Why Use NASA Evidence?
A candidate can sound clear while repeating a misconception. NASA gives you an external reference for judging the central explanation.
This keeps writing style from hiding a factual weakness. Your accuracy score now rests on evidence that another evaluator can inspect.
- Return to the AI Trainer Evaluation Workbook from earlier.
- Select the Evaluations tab.
- Add an Evidence Source label below the constrained-prompt scorecard.
- Paste https://spaceplace.nasa.gov/seasons/ beside the label.
The scorecard now points to the source that supports your factual judgment.
- Add a NASA-Supported Claims label below the source.
- Record Earth's tilted axis causes the seasons as the first supported claim.
The first claim gives you a direct reference for checking each candidate's explanation of axial tilt.
- Record the changing Earth-Sun distance is not the cause as the second supported claim.
- Check that the evidence note contains the NASA link plus both supported claims.
Good work. Your evaluation now has an external factual baseline.
Score both candidates consistently
A pairwise comparison becomes auditable when each response receives a separate score for each requirement. Judge one criterion at a time so a polished writing style does not influence every score.
How to Apply Each Criterion
- Use Factual Accuracy to compare each candidate's central claims with the NASA evidence.
- Use Instruction Following to check the requested audience, the 100-word limit, axial tilt, the distance misconception, plus the source link.
- Use Relevance to check whether each sentence supports the requested explanation.
- Use Clarity for the Audience to judge whether a 12-year-old could follow the explanation.
- Use Source Quality to judge whether the included link supports the factual claims.
- Select the Rubric tab.
- Review the scoring anchors for Factual Accuracy.
- Review the scoring anchors for the other four criteria.
- Return to the Evaluations tab.
- Read Candidate A's constrained-prompt response.
- Enter Candidate A's five scores in the blank scorecard.
- Read Candidate B's constrained-prompt response.
- Enter Candidate B's five scores in the blank scorecard.
Strong progress. Both candidates now have criterion-level scores that another evaluator can trace back to the rubric.
- Calculate Candidate A's total by adding its five scores manually.
- Record Candidate A's total beside its scores.
- Calculate Candidate B's total by adding its five scores manually.
- Record Candidate B's total beside its scores.
- Compare the two recorded totals.
- Record the matching overall outcome as A, B, or Tie.
- Judge your confidence from the strength of the evidence supporting that outcome.
- Record the confidence label as High, Medium, or Low.
You should now see ten criterion scores beside the candidate responses. You should also see two totals plus one overall outcome.
Evidence Determines the Winner
Candidate responses vary between runs. Your result is valid when the evidence trail supports the recorded outcome.
Write and audit the adjudication
An adjudication explains how the scores support the final decision. Three focused sentences make your reasoning quick to inspect.
- Add an Adjudication Note field below the completed scorecard.
- Write the first sentence to name Candidate A, Candidate B, or Tie as the outcome.
Your opening sentence now tells a reviewer exactly what you decided.
- Write the second sentence to cite the strongest criterion-level evidence.
- Write the third sentence to explain how the constrained prompt or rubric made the decision more defensible than the instinctive comparison.
The finished note connects the outcome to a scored criterion. It also captures what changed after you replaced instinct with explicit requirements.
Before you audit the workbook, do you expect the totals plus the written decision to tell the same story?
- Recalculate Candidate A's five scores from top to bottom.
- Compare Candidate A's sum with its recorded total.
- Recalculate Candidate B's five scores from top to bottom.
- Compare Candidate B's sum with its recorded total.
- Compare the recorded A, B, or Tie outcome with the two totals.
Both recalculated sums should match the displayed totals. The selected outcome should follow those totals.
- Confirm that the confidence label is High, Medium, or Low.
- Count the three sentences in the adjudication note.
- Confirm that the NASA evidence link appears beside the scorecard.
- Confirm that the evidence note records both supported claims.
What Should I See?
You'll see ten criterion scores plus two matching totals. Each score will use the existing 1-to-3 scale.
The overall outcome will agree with the totals. The confidence field will contain one allowed label.
The NASA link will support both recorded claims. The three-sentence adjudication will explain the decision.
You did the difficult work of turning judgment into an auditable decision. Next, you'll package this evaluation as a privacy-safe work sample for a reviewer.
Package Your Work Sample
Your evidence-backed AI evaluation is complete. The workbook now supports a defensible decision.
Raw grading demonstrates effort. A reviewer needs a concise Google Sheets artifact that protects private information.
In this step, get ready to:
- Build a concise reviewer summary.
- Sanitize the workbook.
- Share the workbook with Viewer access.
Write the reviewer summary
The Portfolio Summary tab is the reviewer's entry point. It should explain your method without exposing product identities.
- Switch back to the AI Trainer Evaluation Workbook from earlier.
- Select the Portfolio Summary tab.
- Add a Task row that summarizes the comparison of two anonymized responses.
- Add a Rubric Rationale row that explains why separate scoring criteria made the evaluation auditable.
You should now see two short rows that explain the evaluation task plus its scoring method.
- Add a Winner row that repeats the A, B, or Tie outcome from Evaluations.
- Add a Strongest Evidence row containing the decisive criterion plus its NASA-backed support.
The summary now connects your decision to criterion-level evidence.
- Add a Calibration Lesson row that explains how clearer instructions made your decision more defensible.
- Review all five rows for references to Candidate A or Candidate B only.
Good work. Your summary now gives a reviewer the task, method, outcome, evidence, plus the lesson in one place.
Sanitize the workbook
A hiring sample needs a firm privacy boundary. This pass keeps the work relevant to the evaluation.
You can clean this safely because the scored anonymized comparisons stay in place.
- Review the Rubric tab for private text.
- Review the Evaluations tab for private text.
- Remove personal notes, employer material, account details, or other sensitive content from the workbook.
You should now see evaluation material only. The workbook no longer exposes unrelated private information.
- Confirm both seasons comparisons identify their outputs as Candidate A or Candidate B.
- Switch back to the NASA Space Place page from earlier.
- Compare the adjudication's factual claims with NASA's explanation of Earth's seasons.
- Return to the workbook.
- Confirm every factual claim in the adjudication points to the NASA source.
Your evidence trail now leads from the adjudication to an authoritative source. Candidate identities remain anonymous throughout the workbook.
Publish a tested view-only link
Viewer access lets a reviewer inspect the workbook without changing its cells. The final test proves that the sharing settings work outside your signed-in session.
Limit Identity Disclosure
Google notes that the file owner's name can be visible when a link is shared. The owner's email can also be visible.
The specific reviewer route limits that disclosure to one chosen recipient.
- Choose the sharing route that matches your privacy needs below.
Anyone with the link
This route suits a work sample where broader link access is acceptable.
- Click Share in the top-right corner of the workbook.
- Under General access, select Anyone with the link.
- Select Viewer as the access role.
- Click Copy link.
- Click Done.
- Record the copied URL here: your view-only workbook link.
Specific reviewer
Direct sharing restricts access to the Google Account you choose.
- Click Share in the top-right corner of the workbook.
- Enter the reviewer's email address in the sharing field.
- Select Viewer as the access role.
- Click Send.
- Click Share again.
- Click Copy link.
- Click Done.
- Record the copied URL here: your view-only workbook link.
Before you test the link, consider whether the workbook will open without editing controls.
For direct sharing, the Google Account listed as the reviewer performs this check.
- In Google Chrome, click More in the top-right corner.
- Select New Incognito window.
- Sign in with the recipient account if you chose direct sharing.
- Paste your view-only workbook link into the address bar.
- Press Enter.
You'll see the sanitized workbook open through the selected access route. Cell editing is unavailable.
- Select the Portfolio Summary tab.
- Check that all five summary fields are visible.
- Try to edit a blank cell.
The cell remains unchanged. This confirms that the copied link provides Viewer access.
Link Not Opening with Viewer Access?
- Reopen Share to confirm that Anyone with the link uses the Viewer role.
- Confirm that the signed-in Google Account matches the recipient used for direct sharing.
- Copy the link again after correcting the access setting.
Get help with troubleshooting your Google Sheets Viewer link.
You have turned a scored workbook into a review-ready work sample with a tested access path.
Secret mission
Adjudicate a Missing-Evidence Edge Case
Add a test case where the required source passage is missing. You will define an Insufficient Evidence rule that stops evaluators from forcing an unsupported winner.
Clean Up Your Resources
Clean Up Your Resources
This project's free access means the workbook and saved chats create no ongoing project charge. Choose the option below that matches what you want to retain.
Resources you used:
- The sanitized AI Trainer Evaluation Workbook in Google Sheets.
- The tested view-only sharing link for the workbook.
- The saved ChatGPT practice chats.
- Any saved Claude practice chats.
Keep everything running
No action is needed. Choose this option while you are using the workbook as a work sample.
- Keep AI Trainer Evaluation Workbook in Google Drive.
- Keep the tested view-only link for reviewer access.
- Retain the practice chats if you want to revisit the original model responses.
Pause - I'll come back to this later
Close the active browser tabs to stop working for now. Your saved workbook and chats remain available in their services.
- Close the Google Sheets tab.
- Close every ChatGPT practice-chat tab.
- Close the Claude practice-chat tabs if you used Claude.
Delete - I don't want to use this again
Deletion can feel drastic, so the steps below separate access removal from content removal. This option removes the workbook, its sharing access, and the saved practice chats.
- Remove Anyone with the link access under General access if you shared the workbook publicly.
- Remove the specific reviewer's Viewer access from the workbook's sharing controls if you used direct sharing.
- Delete AI Trainer Evaluation Workbook from Google Drive through the workbook's file controls.
- Delete each saved ChatGPT practice chat through that conversation's account controls.
- Delete each saved Claude practice chat through that conversation's account controls.
- Open the copied sharing link in a private browser window to confirm that the workbook is unavailable.
Nice Work!
Nice Work!
You did it! Your sanitized view-only Google Sheets workbook now demonstrates repeatable AI evaluation judgment.
You've learned how to:
- Recorded anonymized model-response pairs through a side-by-side pairwise comparison that made the limits of instinctive judgment visible.
- Created a five-criterion rubric with 1-to-3 anchors for consistent scoring that another evaluator can audit.
- Packaged your source-backed adjudication as a privacy-safe hiring work sample with Viewer access.
- Secret Mission: Added a missing-evidence edge case that returns Insufficient Evidence when the required source passage is absent.
Ready to quiz yourself?