Check the numbers, calculations, and sources in an AI report sentence by sentence
Split a plausible summary into facts, calculations, and interpretations. Recalculate ratios and averages from synthetic monthly data, and rewrite sentences with missing sources or overstated causes into checkable statements.
Content checked 2026.09.19Example files included
Show contents
This translation was generated by AI. Check the code, units, and numbers against the original. Native-speaker review has not yet been completed for each language. 한국어
Who this is forWriters and reviewers who check work reports or summaries prepared by AI
What you need
Obtain the draft report and the source data it used, separately.
Check the units, period, scope, and extraction date of the source data. Mark anything you cannot access as unverifiable.
Prepare a calculator or local calculation tool and create a file to keep the review record.
01Check the evidence, not the confidence of the sentence
Even when an AI report reads smoothly, you still have to check separately whether the numbers and sources are right. NIST's guidance on generative AI risks explains that not only incorrect content but also reasoning and citations can be fabricated, and it recommends reviewing the provenance and citations of outputs. This article separates each sentence that contains a number and links it to the source data, the formula, and the final wording.
02Split sentences into three types
Type
Example
Evidence to check
Fact from source data
1,500 requests were completed in August
The month, column, and definition in the source data
Calculated result
Completed requests rose 25 %
Previous and current month values, denominator, formula
Interpretation or recommendation
AI made us faster, so we should expand it
Data that isolates the cause, and decision criteria
If one sentence mixes these three types, split it into short sentences. In “requests increased and automation was effective,” the first part can be checked by counting, but the second part needs separate evidence. Correct numbers do not make the causal explanation correct.
03Order for checking each sentence
Copy the draft under review and number every sentence that contains a number, ratio, date, or source.
Keep only one claim to check per sentence and move it to claim-log.txt.
Write down the location and definition of the source data. If there is only a URL, open the document yourself and find the table, page, reference date, and figure.
Separate directly quoted values from calculated values. For calculated values, record the numerator, denominator, unit, and rounding rule.
Recalculate totals, changes, ratios, and averages from the source data. Do not end the review by simply asking the same AI once more whether it is right.
Assign one status: verified, needs correction, insufficient evidence, or unverifiable. For sentences with insufficient evidence, narrow the claim instead of keeping it with different numbers.
Keep the revised version together with the review record, and when the source data changes, recheck the affected sentences first.
04Synthetic source data created for this article
Period
Completed
Completed on time
Work hours
Cost
July 2026
1,200
1,080
450 hours
KRW 3,600,000
August 2026
1,500
1,380
525 hours
KRW 4,200,000
Work hours and cost are defined as corresponding to the same set of completed requests. If waiting time, time spent on unfinished requests, or costs from other departments are mixed in, the per-request metrics below no longer mean the same thing. With real data, confirm this correspondence first.
text
검수 연습용 초안 — 오류를 찾기 위해 직접 작성한 문장입니다. 실제 AI 서비스의 출력이 아닙니다.
① 8월 처리 완료 건수는 전월보다 25 % 증가했다.
② 기한 내 완료율은 90 %에서 92 %로 2 % 개선됐다.
③ 두 달 전체 기한 내 완료율은 월별 비율의 평균인 91.0 %다.
④ 건당 평균 작업 시간은 22.5분에서 21.0분으로 6.7 % 줄었다.
⑤ 비용 총액은 증가했지만 건당 비용은 3000원에서 2800원으로 낮아졌다.
⑥ AI 도입 덕분에 작업 시간이 줄었으므로 모든 팀에 바로 확대해야 한다. [근거 자료 없음]
05Recalculate ratios and averages
Item
Formula
Checked value
Growth in completed requests
(1500−1200)÷1200×100
25.0 %
On-time completion rate, July and August
1080÷1200; 1380÷1500
90.0 %; 92.0 %
Difference in on-time completion rate
92.0−90.0
+2.0 %p
Completion rate for both months
(1080+1380)÷(1200+1500)×100
About 91.1 %
Work time per request
450×60÷1200; 525×60÷1500
22.5 min; 21.0 min
Change in time per request
(21.0−22.5)÷22.5×100
About −6.7 %
Cost per request
3600000÷1200; 4200000÷1500
KRW 3,000; KRW 2,800
The two-month rate is calculated by adding the original numerators and denominators, not by simply averaging the monthly rates. Even if 91.0 % and 91.1 % look close, the key point is that they are calculated over different things. The change from 90 % to 92 % is 2 percentage points (pp); as a relative increase it is about 2.2 %.
Work time is given in hours, so multiply by 60 to convert to minutes and then divide by the number of requests. Keep the calculated values at full precision and show one decimal place in the text. Do not recalculate the overall metric from rounded monthly values alone.
06Check sentences with links against the original too
Open the document from the organization's official page or the original data provider. Check that the title, author, and date in the report match the actual document.
Find whether the number used in the claim actually appears in that table or paragraph. Do not verify using only a search-result snippet.
Read the reference period, target population, units, whether it is a sample or an estimate, and whether it has been revised. Do not confuse the publication date with the observation period.
Check whether the draft claims more than the original says. A change in a specific sample does not mean an effect across all teams.
If you cannot access it or there is no supporting paragraph, mark it as unverifiable. Do not fill a missing source with a plausible title or URL.
The numbers in this synthetic example come from source-data.json. The NIST document is only a reference for the review procedure, not a source supporting the fictional monthly figures above. Record the source of the method separately from the source of the numbers.
07Revised sentences and how to check them
text
검수 후 문장 — 합성 자료에 대한 설명이며 실제 성과가 아닙니다.
8월 처리 완료 건수는 1500건으로 7월 1200건보다 25.0 % 증가했다.
기한 내 완료율은 90.0 %에서 92.0 %로 2.0 %p 높아졌다.
두 달 전체 기한 내 완료율은 2460/2700×100으로 약 91.1 %다.
완료 건당 평균 작업 시간은 22.5분에서 21.0분으로 약 6.7 % 감소했다.
총비용은 3600000원에서 4200000원으로 약 16.7 % 증가했고,
완료 건당 비용은 3000원에서 2800원으로 약 6.7 % 감소했다.
자료에는 AI 도입 여부, 작업 난이도, 인력 구성, 다른 변화 요인이 없다.
따라서 차이의 원인이나 확대 도입의 효과는 이 자료만으로 판단할 수 없다.
출처: source-data.json, months[0], months[1]; 설명용 합성 자료.
①, ④, and ⑤ can be kept after checking the definitions and calculations. For ②, change 2 % to 2 pp, and recalculate ③ from the summed numerators and denominators. ⑥ states a cause and a recommendation that are not in the data, so delete it or move it to a separate verification task. Do not record only the incorrect sentences; also record the evidence location for the sentences you kept so they can be rechecked at the next revision.
Recalculated values: 2,700 completed, 2,460 on time, overall rate 91.111… %
Cost check: total cost +16.666… %, cost per request −6.666… %
Scope of conclusion: the difference between the two months can be described, but the causal effect of AI was not established
08Common errors and limits
Do not compare with the same metric after changing the denominator. Received requests and completed requests are different denominators.
Do not fill items missing from the source data with 0 or an average.
An increase in the total and a decrease in a per-request metric can happen at the same time, so do not delete them as contradictory.
Do not approve claims the document does not make just because an official document is linked.
Do not state that checking the arithmetic also verified the truth of the source data, the representativeness of the sample, or legal or financial judgments.
The download includes README.txt, source-data.json, draft-for-review.txt, corrected-summary.txt, claim-log.txt, and expected-checks.json. The numbers were recalculated with local Python. This does not evaluate the accuracy of any real AI service or report generation system. References were checked on the official pages on September 19, 2026.
Execution and verification record
Windows local Python standard library Fraction and Decimal; official sources checked 2026-09-19
Checked totals of completed and on-time requests, monthly and overall rates, and rates of change
Checked hour-to-minute conversion, time and cost per request, and rates of change
Saved the checked values to expected-checks.json and compared them with the values shown in the text
Verification limits
Checked only synthetic demonstration data
Did not call a real AI service, verify real organizational data, or analyze causal effects
Turn “automate this” into an executable task description. Attach a synthetic sample with no sensitive data and a hand-checked expected result to complete a request for a script that totals work logs by team.
Check an aggregation function with four rows you can calculate by hand and 12 unit tests. Verify not only normal values but also empty input, zero, decimals, and invalid input.