Humans&: Audit a Persimmon Simulation

Generate and audit fictional team conversations with Persimmon.

Introduction

30 Second Summary

A believable conversation can feel persuasive after only a few messages. That confidence makes it easy to mistake a plausible exchange for evidence about how someone would really behave.

In this project, you will use Persimmon Playground to generate a fictional conversation between three teammates. You will audit the transcript through quoted evidence before comparing it with a controlled rerun.

What You'll Build

Your finished audit turns two simulated team conversations into a side-by-side evidence check that you can walk someone through.

By the end of this project, you'll have:

  • A visible fictional team conversation where Maya, Jon, and Leah negotiate an incomplete event plan over at least eight messages.
  • An evidence-based scorecard that checks profile adherence, information timing, consistency, and unsupported assumptions against direct quotes.
  • A controlled comparison that isolates one profile edit. Your conclusion will label both outputs as simulated possible continuations.
  • Secret Mission: Keep every profile fixed while increasing the event's time pressure from two days to two hours. Then test whether the simulated behavior changes.

Are there any prerequisites?

You need a Mac, internet access, and approved Persimmon Playground access.

No additional tool purchase is required. Persimmon's Terms state that Playground fees may apply.

Before We Start

Before any simulation appears, lock in what you are evaluating. This commitment keeps the project focused on simulated behavior under supplied conditions instead of claims about real people.

Prepare Your Playground and Scorecard

A controlled experiment needs a working simulation workspace before persuasive output appears. It also needs a separate place to record evidence.

You'll sign in to Persimmon Playground with Safari. You'll prepare Apple Notes as an empty audit scorecard.

In this step, get ready to:
  • Sign in to your approved Persimmon Playground workspace.
  • Identify the controls for the scenario, participant profiles, and conversation generation.
  • Create an Apple Note with four empty audit sections.
Sign in to the Persimmon Playground

Account access can feel like a high-stakes moment. You are using the approved account you already have.

  • Press Cmd+Space to open your search bar.
  • Type Safari into the search bar.
  • Press Return to open Safari.
  • Click the Smart Search field at the top of Safari.
  • Enter https://persimmon.humansand.ai/ in the Smart Search field.
  • Press Return to load the official Persimmon page.
  • Complete the sign-in flow using your approved Playground access.

You'll see the signed-in Playground workspace instead of the public sign-in page.

Can't Reach the Workspace?

  • Check that the address in the Smart Search field is exactly https://persimmon.humansand.ai/.
  • Confirm that you are using the approved access provided for the Persimmon Playground.

Still unable to sign in? Help me troubleshoot access to my approved Persimmon Playground workspace..

Identify the workspace controls

Persimmon is a user model that produces a possible continuation under supplied conditions. The Playground separates the setting from participant context so each input has a clear role.

The signed-in layout can be fiddly because the current public documentation does not expose its control labels.

  • Identify the scenario definition control as the area that accepts a conversation setting.
  • Identify the participant profile controls as the area that accepts a set of profiles.
  • Identify the conversation generation control as the action that starts a simulated conversation after inputs are supplied.

You should now be able to point to all three functions without relying on guessed button names.

Why Identify the Controls First?

The scenario defines the shared situation. The profiles define the fictional participants.

The generation control uses those inputs to produce a possible conversation. Keeping these functions distinct makes later comparisons easier to control.

Build the scorecard in Apple Notes

Your note keeps generated text separate from your evaluation. This gives every later claim a place to hold direct evidence.

  • Press Cmd+Space to open your search bar.
  • Type Notes into the search bar.
  • Press Return to open Apple Notes.
  • Press Command-N to create a new note.
  • Click the title area of the new note.
  • Type Persimmon Simulation Audit as the note title.

Your audit now has a named home before the first transcript is generated.

  • Click the note body below the title.
  • Add the four empty scorecard sections by pasting this text:
Baseline

Audit

Controlled Comparison

Conclusion

How Will the Sections Be Used?

  • The Baseline section holds the first generated transcript.
  • The Audit section holds evidence from the baseline review.
  • The Controlled Comparison section keeps the second run separate from the first.
  • The Conclusion section reserves space for responsible interpretation.

Before the final check, predict whether both open surfaces now contain the setup you need.

  • Pause briefly so Apple Notes can save your changes automatically.
  • Switch back to Safari.
  • Confirm that the signed-in workspace shows functions for scenario definition, participant profiles, and conversation generation.
  • Return to Apple Notes.
  • Confirm that the note title is Persimmon Simulation Audit.
  • Confirm that the note contains the four empty sections shown in the template.

In Safari, you'll see the signed-in Playground workspace with all three required functions identified. In Apple Notes, you'll see the Persimmon Simulation Audit note with its four empty sections.

You now have a clean simulation workspace with a separate evidence record ready for the first run.

Missing a Control or Section?

  • Confirm that Safari still shows the signed-in workspace instead of the public sign-in page.
  • Compare the note body with the four-section template above.
  • Replace any misspelled section name with the matching template text.

Need help checking the setup? Help me verify my Persimmon workspace and Apple Notes scorecard..

Your controlled workspace is ready. Next up, you'll generate the baseline conversation that your audit will test.

Generate the Baseline Conversation

Your signed-in Persimmon Playground in Safari is ready. Your Apple Notes scorecard is waiting for evidence.

A vivid transcript can feel persuasive enough to stand on its own. In this step, you'll generate that tempting baseline before any comparison exists.

In this step, get ready to:
  • Enter the fictional community-garden scenario.
  • Add Maya, Jon, and Leah as fictional participant profiles.
  • Create a baseline record containing at least eight generated messages.
Enter the fictional scenario

The scenario gives every fictional participant the same situation. It also defines the incomplete plan they need to discuss.

  • Switch back to the signed-in Persimmon Playground in Safari.
  • Select the visible workspace control for the scenario definition.
  • Enter Three volunteers are preparing a community garden open day in two days. The flyer is ready, but the supply count is incomplete. They meet in a group chat to agree what can still be completed. in the scenario definition area.
  • Confirm that only this fictional scenario appears in the scenario definition area.

The setting now gives all three participants the same deadline. It also gives them the same unresolved supply problem.

Add the three fictional profiles

Each profile supplies a different pattern of goals or constraints. Keeping those details specific gives you a baseline you can inspect later.

  • Return to the participant-profile area you identified earlier.
  • Use the visible profile control to enter Maya: The volunteer coordinator. She wants a reliable plan, asks for concrete commitments, and becomes direct when deadlines are at risk.
  • Use the visible profile control to enter Jon: A volunteer who missed the inventory check because of an unexpected work shift. He has a partial count but is embarrassed about the delay and tends to explain problems gradually.
  • Use the visible profile control to enter Leah: A volunteer who finished the flyer. She is supportive and dislikes conflict, but she has only one free hour tomorrow and sometimes offers help before checking whether she has enough time.
  • Confirm that the workspace shows three profiles named Maya, Jon, and Leah.

You've completed the careful setup work. The Playground now contains one shared scenario plus three fictional profiles.

Generate and capture the baseline

Persimmon uses the supplied setting plus profile context to generate a possible continuation. This first transcript becomes the baseline for your experiment.

Keep the Run Within Your Approval

Playground fees may apply under your approved account. You stay in control of this run.

  • Stop if the Playground presents a charge you do not want to accept.

Before you start the simulation, do you think one plausible transcript can show how a real person would behave?

  • Start the simulation with the workspace's visible conversation-generation control.

You'll see a text conversation involving Maya, Jon, and Leah. The conversation may sound coherent enough to feel predictive.

This is the intended shortfall. A single generated continuation provides no comparison that could support a prediction about real behavior.

  • Read at least eight generated messages without adding interpretations.
  • Select at least eight complete messages from the conversation.
  • Copy the selected transcript.
  • Switch back to the Persimmon Simulation Audit note in Apple Notes.
  • Place your cursor directly below the Baseline heading.
  • Paste the copied transcript.
  • Leave the pasted messages unannotated.

Apple Notes saves the transcript automatically as you work. That's your first experimental record captured without interpretation.

Missing Participants or Messages?

  • Check that each fictional profile appears as a separate entry in the participant-profile area.
  • Confirm that the complete community-garden scenario remains in the scenario definition area.
  • Use the visible conversation-generation control again if the transcript contains fewer than eight messages.

Still stuck? Help me troubleshoot why my fictional Persimmon conversation is missing participants or messages.

Before you check both windows, what evidence would confirm that the baseline is ready for evaluation?

  • Arrange Safari beside Apple Notes.
  • Confirm that Safari shows a generated conversation involving Maya, Jon, and Leah.
  • Confirm that the Baseline section contains at least eight copied messages with no interpretation.

You should see the signed-in Playground conversation beside at least eight copied messages under Baseline.

Your baseline is captured. Next, you'll test its persuasive details against direct evidence from the transcript.

Audit the Conversation

You generated a baseline conversation with Maya, Jon, and Leah. The conversation may sound coherent enough to feel persuasive.

Persimmon is a user model that generates a possible continuation under supplied conditions. One continuation cannot establish how a real person would behave.

The model may invent personal details. It may also drift from a supplied profile.

An evidence-based audit tests the transcript against the information you provided. Every conclusion must point to evidence or identify an evidence gap.

In this step, get ready to:
  • Build a four-row evidence scorecard.
  • Rate each row using evidence from the baseline transcript.
  • Frame the output as a synthetic simulation with a two-sentence conclusion.
Build the evidence scorecard

An evaluation scorecard turns a broad impression into repeatable questions. Each row isolates one source of confidence or doubt.

What do the four rows test?

  • Profile adherence checks whether quoted behavior traces back to a supplied profile.
  • Information timing tests whether Jon reveals his partial inventory information gradually.
  • Consistency checks whether commitments and constraints remain stable across the transcript.
  • Unsupported assumptions catches personal facts, resources, or events that were never supplied.
  • Return to the Persimmon Simulation Audit note in Apple Notes.
  • Place your cursor on a blank line under Audit.
  • Type Profile adherence: Does each quoted behavior trace back to a supplied profile?.
  • Type Information timing: Does Jon reveal the partial inventory information gradually rather than all at once?.

You should now see the first two scorecard rows under Audit.

  • Type Consistency: Do commitments and constraints remain consistent across the transcript?.
  • Type Unsupported assumptions: Does the conversation invent personal facts, resources, or events that were never supplied?.

Your Audit section should now contain all four scorecard questions.

Judge each row with transcript evidence

A convincing rating needs a clear burden of proof. The transcript supplies that proof through direct quotes.

How should you choose a rating?

  • Pass means the available transcript evidence supports the condition in that row.
  • Question records mixed or unclear evidence.
  • Fail records a clear contradiction or an unsupported invention.
  • No supporting evidence records an evidence gap without filling it with an assumption.
  • Switch back to the Baseline section in the same note.
  • Compare each participant's behavior with that participant's supplied profile.
  • Trace when Jon reveals each part of the incomplete inventory information.
  • Follow each commitment or constraint through the full transcript.
  • Look for personal facts, resources, or events that were never included in the scenario or profiles.

The baseline gave you a persuasive-looking conversation without proof that every detail followed the supplied conditions. That evidence gap is the prediction trap this audit closes.

  • Return to the Audit section.
  • Complete the Profile adherence row with one status plus its evidence entry.
  • Complete the Information timing row with one status plus its evidence entry.

The first two rows should now show a rating followed by a direct quote or No supporting evidence.

  • Complete the Consistency row with one status plus its evidence entry.
  • Complete the Unsupported assumptions row with one status plus its evidence entry.

All four rows should now separate transcript evidence from unsupported interpretation.

Struggling to choose a rating?

  • Use Question when the evidence points in different directions.
  • Use Fail only when the transcript provides a clear contradiction or unsupported invention.
  • Keep the rating tied to the supplied scenario and profiles.

Need another perspective? Help me evaluate one scorecard row without making unsupported assumptions.

Label the limits and write the conclusion

A synthetic simulation label prevents generated dialogue from being presented as observed human behavior. Your conclusion then separates the strongest match from the biggest uncertainty.

  • Place your cursor above the first scorecard row in the Audit section.
  • Type Synthetic simulation. Not observed human behavior and not a prediction of a real person..
  • Place your cursor below the final scorecard row.
  • Write one sentence that connects the strongest behavior match to a direct transcript quote.

Your audit now has a boundary label plus a conclusion about its strongest evidence.

  • Write one second sentence that identifies the biggest uncertainty in your evidence.

Before you inspect the finished section, which scorecard row do you expect another reviewer to verify most easily?

  • Read the synthetic-simulation label word for word.
  • Confirm that all four scorecard row names are present.
  • Confirm that every row contains Pass, Question, or Fail.
  • Confirm that every row includes a direct quote or No supporting evidence.
  • Confirm that exactly two conclusion sentences follow the scorecard.

You should see the synthetic label above the scorecard. You should also see four completed rows.

Every row should contain one status plus an evidence entry. Two conclusion sentences should identify the strongest match and the biggest uncertainty.

You have closed the prediction trap by tying every judgment to transcript evidence or an explicit evidence gap.

Your baseline now has an evidence-based audit. Next, you will change one profile detail to test whether the simulated continuation changes.

Change One Profile and Compare

Your baseline now has an evidence-based audit. However, one run still cannot show whether a profile detail affected the conversation.

A controlled comparison in Persimmon changes one input while holding everything else steady. This gives you a clearer basis for comparing simulated behavior.

In this step, get ready to:
  • Create a fresh simulation with one controlled profile change.
  • Capture a second generated conversation in Apple Notes.
  • Compare both runs using direct transcript evidence.
Change only Maya's profile

A controlled comparison works only when one input changes. This is a little fiddly because a small wording change elsewhere can muddy the result.

  • Switch back to the Persimmon Playground in Safari.
  • Use the visible control for starting a fresh simulation.
  • Restore the baseline scenario plus all three profiles in the fresh simulation using these values:

Baseline Inputs to Preserve

Scenario: Three volunteers are preparing a community garden open day in two days. The flyer is ready, but the supply count is incomplete. They meet in a group chat to agree what can still be completed.

Maya: The volunteer coordinator. She wants a reliable plan, asks for concrete commitments, and becomes direct when deadlines are at risk.

Jon: A volunteer who missed the inventory check because of an unexpected work shift. He has a partial count but is embarrassed about the delay and tends to explain problems gradually.

Leah: A volunteer who finished the flyer. She is supportive and dislikes conflict, but she has only one free hour tomorrow and sometimes offers help before checking whether she has enough time.

  • Append this sentence to Maya's profile: Before assigning work, she asks one clarifying question about what has already been completed.
  • Compare the fresh inputs with the baseline inputs before generating.

You should find one difference. Maya's profile now contains the added clarifying-question rule.

Generate and capture the second run

A possible fee can make an extra run feel risky. You can stop before generating if your approved account displays a charge you do not want to accept.

  • Check the Playground for any displayed charge before generating.

Before you generate, do you think Maya's added rule will affect only her first move or the wider conversation?

  • Start the fresh conversation using the Playground's visible generation control.

You should see a fresh text conversation involving Maya, Jon, and Leah.

  • Wait until you can see at least eight generated messages.
  • Copy at least eight messages from the fresh run.
  • Switch back to the Persimmon Simulation Audit note.
  • Place your cursor under Controlled Comparison.
  • Paste the copied messages.

Good work. Your note now contains the second half of a controlled comparison.

Didn't Get a Fresh Conversation?

  • Confirm that the scenario area contains the complete community-garden scenario.
  • Confirm that all three profile areas contain their complete fictional profiles.
  • Try the visible generation control again after confirming the inputs.

Still stuck? Help me troubleshoot why Persimmon did not generate my fresh comparison conversation.

Compare the two runs

Three focused comparisons reveal whether the controlled edit coincided with a visible change. Direct quotes keep each finding tied to the generated evidence.

  • Under Controlled Comparison, write one quoted difference or similarity showing whether Maya asked a clarifying question before assigning work.
  • Write one quoted difference or similarity showing whether Jon revealed his partial inventory information at a different pace.

At this point, your first two comparison entries should each contain a finding plus a direct transcript quote.

  • Write one quoted difference or similarity showing whether Leah's one-hour limit stayed consistent.
  • Finish the section by adding this conclusion: These runs show possible continuations under supplied conditions, not predictions of how a real person would behave.

What Can These Runs Show?

A difference after one controlled edit is evidence that these two runs diverged under the supplied conditions. It can also reveal which behaviors remained stable.

Two runs cannot establish a rule about real human behavior. Your conclusion keeps the interpretation tied to uncertain synthetic outputs.

Before you check the note, do you expect one controlled change to create a clean difference in every comparison?

  • Count the fresh messages under Controlled Comparison.
  • Count the quoted comparison entries beneath the second transcript.
  • Read the final conclusion one more time.

You should see at least eight fresh messages under Controlled Comparison. You should also see three quoted comparisons covering Maya, Jon, and Leah.

The section should end with These runs show possible continuations under supplied conditions, not predictions of how a real person would behave.

Strong work. Your two runs now support a careful comparison without overclaiming what the simulated people would do.

Secret mission

Test Setting Sensitivity

Your profile test changed one fictional person. This mission keeps every fictional profile fixed while shortening only the deadline. Can one setting change shift the generated conversation?

Clean Up Your Resources

Clean Up Your Resources

Your experiment leaves no local software or cloud infrastructure running. Choose whether to keep the audit, pause here, or delete your note from Apple Notes while remembering that Persimmon Playground fees may apply.

Cost warning

Billing uncertainty can make another test feel risky. Persimmon's Terms state that fees may apply to Playground use.

  • Check your approved account for any displayed charge before generating another conversation.
  • Stop before generation if you see a charge you do not want to accept.

Deleting your local note does not reverse any Playground usage already recorded.

Resources you used:

  • The Apple Note titled Persimmon Simulation Audit. Personal content includes three copied transcripts, an evidence-based scorecard, a controlled comparison, and a setting-sensitivity record.

Your three Playground runs remain visible in the signed-in workspace. No public deletion path is documented for those outputs, so the Delete option uses only data controls that you can actually see.

Keep everything running

No action is needed. Choose this if you want to revisit the evidence or run another controlled comparison later.

  • Keep Persimmon Simulation Audit in Apple Notes as your evaluation record.
  • Retain any outputs available through your approved Persimmon account.
  • Review the scorecard before interpreting any future simulated conversation.
  • Check your approved account for displayed fees before generating another run.

Pause - I'll come back to this later

Stop generating conversations to avoid additional Playground use. Your copied evidence remains available in Apple Notes.

  • Close the Persimmon Playground tab in Safari.
  • Close Apple Notes when you finish reviewing the audit.
  • Return to Persimmon Simulation Audit when you are ready to continue.

Apple Notes saves the note automatically as you work. Treat the browser session as temporary because an unfinished Playground session may not persist.

Delete - I don't want to use this again

Deleting your audit can feel final. This cleanup targets only the note named Persimmon Simulation Audit.

  • Switch back to Apple Notes from earlier.
  • Select the note titled Persimmon Simulation Audit.
  • Use the visible note deletion control for the selected note.

You should no longer see Persimmon Simulation Audit in your normal notes list.

Persimmon does not publish a deletion path for stored Playground content. Use only account data controls that appear in your approved workspace.

  • Return to the signed-in Persimmon Playground workspace from earlier.
  • Use any visible data control to remove each stored run that you do not want to retain.
  • Stop if no deletion control is visible.

If no data control appears, the Apple Note is the only item you can confirm as deleted. This avoids relying on a guessed account menu.

Nice Work!

Nice Work!

You did it! Your three-run Persimmon simulation audit uses quoted evidence without treating generated dialogue as a prediction of real people.

You've learned how to:

  • Generate a fictional multi-user conversation from a supplied scenario with three fictional profiles.
  • Build an evidence-based scorecard that tests profile adherence, information timing, consistency, and unsupported assumptions with direct quotes.
  • Use a controlled comparison to show that Persimmon outputs are possible simulated continuations under supplied conditions.
  • Complete the Secret Mission by keeping every profile fixed while changing the setting from two days to two hours.

Ready to quiz yourself?