Case Facts
Michael S. Weber died a resident of Saratoga County on December 7, 2003. His sister, Susan F. Weber, was named executrix of his estate and trustee of a trust for the benefit of his son, Owen K. Weber. Petitioner filed for judicial settlement of the trust’s interim account on June 17, 2022. Objectant responded with objections alleging petitioner breached her fiduciary duty in retaining and personally using a piece of real property on Cat Island, the Bahamas, that the trust had received in-kind from the estate. A hearing was held over three days in May and June 2024. To support his damages claim, objectant called materials expert Charles Ranson, who had prepared a “Preliminary Expert Report” dated December 14, 2022, and, mid-hearing, a “Supplemental Damages Report” dated May 28, 2024.
The Setup
Testimony at the hearing revealed that Ranson had relied on Microsoft Copilot - described by the court as “a large language model generative artificial intelligence chatbot” - to cross-check the calculations in his supplemental damages report. Ranson could not recall what input or prompt he had used, could not identify what sources Copilot drew on, and could not explain how the tool arrived at its output. No testimony addressed whether Copilot’s calculations accounted for fund fees or tax consequences. Ranson maintained that using AI tools to help draft expert reports is generally accepted in the fiduciary-services field and represents the future of such work - but he could not name a single publication or source to support that claim.
The Failure
To test the reliability of Copilot’s output for itself, the court ran its own experiment. On a computer issued by the state’s Unified Court System, it entered the calculation Ranson’s analysis depended on: “Can you calculate the value of $250,000 invested in the Vanguard Balanced Index Fund from December 31, 2004 through January 31, 2021?” Copilot returned $949,070.97 - a number different from Ranson’s own figure. Running the identical query on two more court-issued computers produced $948,209.63 and “a little more than $951,000.00,” respectively.
The court then asked Copilot directly whether it was accurate. It answered:
“I aim to be accurate within the data I’ve been trained on and the information I can find for you. That said, my accuracy is only as good as my sources so for critical matters, it’s always wise to verify.”
Asked whether its calculations were reliable enough for use in court, Copilot responded:
“When it comes to legal matters, any calculations or data need to meet strict standards. I can provide accurate info, but it should always be verified by experts and accompanied by professional evaluations before being used in court.”
The Ruling
The court found Ranson’s damages calculations “inherently unreliable” on several independent grounds as well: he started his analysis from an incorrect date, omitted real estate taxes and other trust expenses, supported his market assumptions with a hearsay phone call, and - by his own admission - skipped the full industry-standard analysis he said was too costly to perform. But the court gave Ranson’s AI use its own dedicated section of the decision. It wrote that it had “no objective understanding as to how Copilot works,” since none was offered at the hearing, and that “the record is devoid of any evidence as to the reliability of Microsoft Copilot in general, let alone as it relates to how it was applied here.”
Applying New York’s Frye standard for scientific and expert evidence - which requires a method to be generally accepted in its relevant field - the court held, on what it called an apparent issue of first impression in Surrogate’s Court practice, that before AI-generated evidence is offered, “counsel has an affirmative duty to disclose the use of artificial intelligence,” and that such evidence “should properly be subject to a Frye hearing prior to its admission.” The court defined “Generative Artificial Intelligence” as AI capable of producing new content in response to a prompt by drawing on a large reference database, distinguishing it from merely “assistive” AI use that supports but doesn’t generate a document or analysis outright. Objectant’s objections were denied in their entirety, the interim accounting was approved, and the trustee’s commissions - more than $108,400 - were allowed.
The Kicker
Asked by the court whether it was reliable, Copilot itself answered that it’s “always good to have a second opinion.” The expert whose report depended on it apparently didn’t get one.
How the AI Issue Unfolded
→ Objectant’s damages expert used Microsoft Copilot to cross-check his supplemental report but could not recall his prompts or explain the tool’s methodology.
→ No one disclosed the AI use before it surfaced in hearing testimony.
→ The court ran the same calculation itself on three separate computers and got three different dollar figures.
→ The expert could not cite any source supporting his claim that AI-assisted report drafting is accepted practice in his field.
→ Independent of the AI issue, the expert used the wrong start date, omitted real taxes and expenses, and skipped the industry-standard analysis he called too costly.
→ The court held that AI-generated evidence must be disclosed in advance and, going forward, may be subject to a Frye reliability hearing before admission.
The Lesson
This is one of the first written decisions to set an actual disclosure rule for AI-generated expert evidence rather than simply penalizing an expert after the fact for an undisclosed hallucination. The court didn’t stop at finding one expert’s numbers wrong - it built a framework for the next case: define what counts as generative versus assistive AI, require its use to be disclosed before the evidence is offered, and treat the reliability question as one for a Frye hearing rather than an afterthought raised on cross-examination. The court’s own three-computer experiment is also notable - it didn’t take the unreliability of the tool on faith any more than it took the expert’s numbers on faith; it tested the premise itself.
Takeaways
If you’re retaining experts:
Ask, in writing, whether any AI tool touched the report or the underlying calculations, and get an answer detailed enough to survive cross-examination - what tool, what prompts, what it was checked against. An expert who can’t reconstruct that process afterward has effectively made the calculation unreproducible, which is its own reliability problem apart from anything AI-specific.
If you’re an expert witness:
If you use a generative AI tool at any stage, keep a record of the prompts and outputs, and be able to explain - in your own words, not the tool’s - why the result is trustworthy. Being unable to say what a tool is that you plan to rely on, or why it’s accepted in your field, is not a small gap; it can be the difference between a report that survives a hearing and one that doesn’t.
If you’re challenging an expert’s AI use:
Don’t just argue the tool is unreliable - test it, the way the court did here. Running the expert’s own query and showing the answer changes between runs, or between machines, is more persuasive than a general objection to AI in the abstract.



