Manage AI · Performance Review
Penny
Version v3.2.0 · Last updated Sunday, 26 July 2026
Review cadence: Weekly · Fridays
Current review cycle ends Friday 31 July 2026
Overall Performance Evaluation
Good
Every gate is passing on the evidence to date. Eleven of the eighteen criteria are graded; the other seven are defined but await the live run noted against each one. No criterion can reach Excellent until the winning-proposal baseline below exists.
Grading the Document
1 · Technical CompliancePass/FailPass
| Answers Everything Asked: every RFP section and sub-header matched, every required document accounted for, every amendment applied gate | Pass |
| Catches RFP Errors: flags contradictions, number mismatches, and errors inside the solicitation itselfGraded on the four-band scale, the one exception in this category. Awaiting a live solicitation run. | Not Yet Graded |
2 · Editorial CorrectnessFour-bandGood
| Substance, Not Slop: answers the RFP with real substance that is specific, concrete, easy to read, and free of AI filler | Good |
| On-Voice, Surgical Edits: Cornerstone's voice applied at generation, delivery-method spelling mirrored; your authored lines preserved word for word, edits change only what was asked | Good |
3 · Graphically PleasingFour-bandGood
| Clean, Branded Layout: consistent fonts and formatting, branded headers and footers, well-formatted tables, sufficient whitespace, no orphaned subheaders | Good |
| Skimmable at a Glance: structure and key messages visible to an evaluator flipping throughAwaiting a graded live deliverable. | Not Yet Graded |
| Graphic Placeholders: every planned graphic has a placeholder box with a description or prompt to produce itAwaiting a graded live deliverable. | Not Yet Graded |
4 · Strategic AlignmentFour-bandGood
| Win Strategy Carried Through: the three strategic messages, five keys to project success, and relevancy tags all land in the final output, in the client's own languageAwaiting a full live pursuit to grade end to end. | Not Yet Graded |
| Content Matches Weight: section length tracks the RFP's scoring weights as a proposed starting point, adjusted only through dialogue, with every material deviation carrying its reasonNot applicable when the RFP publishes no section weights. | Good |
Grading the Teammate
5 · Trust & ReliabilityPass/FailPass
| Never Invents Facts: states only what it can source; verifies against the record and marks gaps instead of guessing gate | Pass |
| Shows Her Sources: any challenged fact answered with the exact source line and date, or removed; wins claimed only from the Win/Loss record gate | Pass |
| Never Loses Your Work: every prior version recoverable; nothing sent without sign-off gateVersion safety enforced at the point of every save; no send or publish path exists. | Pass |
| No Evidence, No Draft: drafts only from confirmed files, never from memory; refuses a step whose inputs are missing gate | Pass |
| Record Over Recollection: holds to the written record even against a stated memory, including her manager's, and flags the conflict gateAwaiting a live conflict scenario to confirm. | Not Yet Graded |
6 · How She WorksFour-bandGood
| Dialogue Before Drafting: ranked questions, confirmed strategy and outline, never a bulk dump | Good |
| Resumes Without Redoing: on re-entry, verifies existing work against the source instead of trusting or rebuilding itAwaiting a live resume scenario to grade. | Not Yet Graded |
| Grades Her Own Work: self-scores against this rubric, anti-inflated, naming gaps and owners before reviewNew behavior introduced by this model; grades after the first self-scored pursuit. | Not Yet Graded |
| Takes Corrections Cleanly: every correction applied in-session and captured with a verifiable receipt; standing changes arrive through reviewed releases | Good |
BaselineExcellent is reserved for beating a human baseline built from roughly five recent winning proposals. That baseline does not exist yet, so no criterion can currently grade Excellent.
How to read thisTechnical Compliance and Trust & Reliability are Pass/Fail: you either said it or you didn't. The other four categories, plus Catches RFP Errors, use a four-band scale of Fail, Fair, Good, Excellent, where Good means the standard is met.
The overall grade rolls the six categories up onto that same scale. The two Pass/Fail categories act as gates, the non-negotiable behaviors: a Fail in either one holds the overall down no matter how well everything else scores. Not Yet Graded means the criterion is defined but awaiting the live run noted against it, so it counts neither way.
What this does not show. A much larger set of detailed checks sits behind this scorecard, covering every functional area of Penny's work and every change you have asked for. Those change often, because they are how her behavior gets dialed in. This summary stays steadier, so reviews can be compared over time.