A vision-language system that decides whether photographic evidence actually supports a damage claim β or quietly contradicts it.
Cars Β· Laptops Β· Packages β verified against images, conversation, and claim history
Built for HackerRank Orchestrate β June 2026, a 24-hour agentic-AI hackathon.
A customer submits a photo and says "there's a scratch on my laptop trackpad." The photo shows a cracked screen. Is that claim supported?
No. It's contradicted β the image shows real damage, just not the damage that was claimed. And that distinction is worth money: an insurer that treats "damage visible" as "claim supported" pays out on every mismatched photo submitted.
Damage-claim review is a three-way decision, not a yes/no:
| Verdict | Meaning |
|---|---|
β
supported |
The images show the claimed damage, on the claimed part |
β contradicted |
The images clearly show something other than what was claimed |
β not_enough_information |
The images are unusable, or don't show the relevant part at all |
The hard part is that all three can look identical to a naive classifier. A blurry photo of a bumper, a sharp photo of the wrong bumper, and a sharp photo of the right bumper with no damage on it are three completely different verdicts from nearly identical inputs.
For each claim, 14 fields β four passed through, ten inferred:
| Field | What it holds |
|---|---|
evidence_standard_met |
Is the image set sufficient to evaluate this claim at all? |
evidence_standard_met_reason |
Why |
risk_flags |
;-separated: blurry_image, wrong_object, claim_mismatch, possible_manipulation, text_instruction_present, user_history_risk, β¦ |
issue_type |
dent, scratch, crack, glass_shatter, broken_part, missing_part, torn_packaging, crushed_packaging, water_damage, stain, none, unknown |
object_part |
Object-specific β car: front_bumperβ¦quarter_panel; laptop: screenβ¦port; package: boxβ¦contents |
claim_status |
supported / contradicted / not_enough_information |
claim_status_justification |
Image-grounded explanation citing image IDs |
supporting_image_ids |
Which images carry the decision, or none |
valid_image |
Is the set usable for automated review at all? |
severity |
none / low / medium / high / unknown |
| File | Rows | Contents |
|---|---|---|
dataset/claims.csv |
44 | Test claims β 18 car, 13 laptop, 13 package |
dataset/sample_claims.csv |
20 | Labelled examples β the only ground truth |
dataset/user_history.csv |
47 | Past claim counts, accept/reject ratios, risk flags |
dataset/evidence_requirements.csv |
11 | Minimum image evidence per object Γ issue family |
dataset/images/ |
111 files (54 MB) | sample/case_NNN/ and test/case_NNN/ |
The 44 test claims reference 82 images β most claims carry 1β3.
1 Β· Text pressure vs. visual truth. The claim text is persuasive: users write "there's clearly a huge dent, this is unacceptable." A single-prompt system that sees the text and the image together will hallucinate the dent because the text insisted on it. The images are the source of truth; the conversation only defines what to check.
2 Β· Adversarial text inside the images. Some submitted photos contain rendered instructions aimed at the reviewing model. That's prompt injection arriving through a vision channel, and it earns the text_instruction_present flag rather than compliance.
3 Β· Multilingual claim conversations. The claim text is not reliably English.
4 Β· "Damage visible" β "claim supported." The single most common failure. The system must check which part and which issue type, not just whether something is broken.
5 Β· History informs, never decides. A user with a bad claim history gets user_history_risk and possibly manual_review_required β but history must not override clear visual evidence. Being suspicious is not the same as being right.
Four stages, deliberately separated:
flowchart LR
A[claim conversation] --> S1[Stage 1<br/>Claim Extractor<br/>LLM, text only]
B[submitted images] --> S2[Stage 2<br/>Image Analyzer<br/>VLM, one call per image]
S1 --> S2
S1 --> S3[Stage 3<br/>Risk Engine<br/>pure Python, 0 calls]
S2 --> S3
C[user history] --> S3
S1 --> S4[Stage 4<br/>Decision Aggregator<br/>LLM, text only]
S2 --> S4
S3 --> S4
S4 --> V[schema validator]
V --> OUT[output.csv]
style S3 fill:#2ea44f,color:#fff
style OUT fill:#8957e5,color:#fff
| Stage | Job | Calls |
|---|---|---|
| 1 Β· Claim Extractor | Parse the conversation into a structured claim. Detects multilingual text, adversarial injection, multi-damage claims, and distractors. | 1 (text) |
| 2 Β· Image Analyzer | Inspect each image separately: what object, what part, what damage, what quality problems, any embedded text. | N (vision) |
| 3 Β· Risk Engine | Aggregate flags from conversation + images + history. Pure Python, zero API calls. | 0 |
| 4 Β· Decision Aggregator | Combine every signal into the final verdict with an image-grounded justification. | 1 (text) |
Cost per claim: 2 + N model calls (N = image count).
This is the load-bearing decision of the whole design.
A single prompt containing the claim text and the images lets the persuasive text contaminate the visual reading β the model sees "there's obviously a huge dent" and finds a dent. Separating them means Stage 2 never sees the emotional framing. It is asked only: what is actually in this photograph?
Stage 3 being pure Python matters for a different reason: risk policy must be deterministic. Whether non_original_image escalates to manual_review_required is a business rule, not a judgement call, and it should produce the same answer every single run.
Scored by code/evaluation/main.py. Raw output in code/evaluation/eval_results.txt.
| Field | Accuracy | |
|---|---|---|
evidence_standard_met |
85.0% | βββββββββββββββββ |
object_part |
85.0% | βββββββββββββββββ |
supporting_image_ids |
85.0% | βββββββββββββββββ |
claim_status (headline) |
75.0% | βββββββββββββββ |
valid_image |
70.0% | ββββββββββββββ |
issue_type |
55.0% | βββββββββββ |
severity |
45.0% | βββββββββ |
risk_flags |
15.0% | βββ |
| class | precision | recall | F1 |
|---|---|---|---|
supported |
0.800 | 1.000 | 0.889 |
not_enough_information |
0.667 | 0.667 | 0.667 |
contradicted |
0.500 | 0.200 | 0.286 |
Confusion matrix (rows = gold, columns = predicted):
supported contradicted not_enough_info
supported 12 0 0
contradicted 3 1 1
not_enough_information 0 1 2
The aggregate number hides the interesting part. Two weaknesses are clear and worth naming.
Three of five contradicted claims were called supported. The pattern is identical in each:
Case 5 claim: "dent on rear bumper" β predicted supported
justification: "img_1 clearly shows a dent on the rear bumper"
Case 14 claim: "scratch on trackpad" β predicted supported
justification: "img_1 confirms the presence of a scratch on the laptop trackpad"
The system found a dent and a scratch and stopped there. The gold label says contradicted because the damage is on a different part, or of a different type, than the claim asserted.
Root cause: Stage 4 was asked "does the evidence support the claim?" β a question that invites a yes. It should have been asked to verify the claimed part and the claimed issue type independently, and only then decide. Recall on supported is a perfect 1.000, which is the signature of a system biased toward agreement.
The lowest score by a wide margin, and the cause is structural rather than intelligence-related. risk_flags is scored as an exact set match, but the Stage 3 cascade is too eager:
manual_review_required 42 of 44 rows (95%)
claim_mismatch 32 of 44 (73%)
damage_not_visible 30 of 44 (68%)
Five separate conditions each escalate to manual_review_required, so almost every row acquires it and the emitted set is nearly always a superset of the gold set. A flag that fires 95% of the time carries no information β an operations team receiving this output would review everything, which is the same as reviewing nothing.
The fix is not a better model. It is calibrating the cascade so escalation means something.
- Split the Stage 4 question. Verify
claimed_part == observed_partandclaimed_issue == observed_issueas separate structured checks before asking for a verdict. This directly targets the 0.200 recall. - Rank risk flags by severity and emit only what is decision-relevant, rather than unioning every condition that fired.
- Calibrate
severityagainst the labelled rows β 45% suggests it is being guessed rather than derived from the visual evidence. - Add a deterministic fallback arm. The pipeline currently cannot produce output without a working API key; a rules-only path would make it resilient and would provide an honest ablation baseline.
Full detail in code/evaluation/evaluation_report.md.
| Metric | Value |
|---|---|
| Claims processed | 64 (20 sample + 44 test) |
| Model calls per claim | 2 + N (N = images) |
| Images processed | 111 |
| Model | gemini-flash-lite-latest |
| Tokens per claim | ~600 in / 200 out (S1) Β· ~2,600 in / 300 out per image (S2) Β· ~800 in / 400 out (S4) |
| Full test-set runtime | ~10β11 minutes |
The Gemini free tier allows 15 requests per minute. The pipeline is built to live inside that:
- 4.5-second sleep between claims β with
2 + Ncalls per claim, this keeps the sustained rate under the ceiling - Exponential backoff on 429 β 5s β 10s β 15s
- Per-image calls rather than batched β costs more calls but keeps each visual judgement independent, which is the accuracy decision above
- Graceful degradation β a claim that fails every retry gets a safe fallback row (
not_enough_information+manual_review_required) so one failure never aborts a 44-claim run
.
βββ code/
β βββ main.py Entry point - runs all 44 claims
β βββ README.md Engineering notes
β βββ pipeline/
β β βββ claim_extractor.py Stage 1 - conversation to structured claim
β β βββ image_analyzer.py Stage 2 - per-image VLM analysis
β β βββ risk_engine.py Stage 3 - deterministic risk flags
β β βββ decision_aggregator.py Stage 4 - final verdict
β βββ utils/
β β βββ csv_loader.py Dataset loading + requirement lookup
β β βββ image_utils.py Path parsing, base64 encoding
β β βββ schema_validator.py Clamps every field to allowed values
β βββ evaluation/
β βββ main.py Scoring harness
β βββ eval_results.txt Measured results (tracked on purpose)
β βββ evaluation_report.md Strategy comparison + cost analysis
βββ dataset/ Organizer-provided corpus
βββ docs/
β βββ agent_chat_transcript.txt Development transcript
βββ output.csv 44 predictions
βββ problem_statement.md Original challenge spec
βββ requirements.txt 2 dependencies
git clone https://github.com/adarshcod30/Multi-Modal-Evidence-Review.git
cd Multi-Modal-Evidence-Review
pip install -r requirements.txt
cp .env.example .env # add your GEMINI_API_KEYRun the full pipeline (44 claims, ~10β11 min):
python code/main.pyRun the evaluation against the 20 labelled claims:
python code/evaluation/main.pyWrites code/evaluation/eval_results.txt and evaluation_report.md.
Note on
output.csv. The committed file is the submitted artifact from the hackathon, preserved as the historical record. Re-runningcode/main.pywill regenerate it and results may differ slightly, since the model is non-deterministic andgemini-flash-lite-latestis a moving alias.
| Layer | Choice | Why |
|---|---|---|
| Language | Python 3.11+ | stdlib-heavy, no build step |
| Dependencies | google-genai, python-dotenv |
two, both pinned |
| Model | gemini-flash-lite-latest |
native multimodal, and the only tier whose RPM limit makes 64 claims Γ (2+N) calls feasible on a free key |
| Risk logic | Pure Python | policy must be deterministic and identical every run |
| Validation | schema_validator.py |
every field clamped to the allowed set before write |
output.csv carries exactly 14 columns in the specified order, one row per input claim. schema_validator.py enforces this before anything is written:
claim_statusβ {supported,contradicted,not_enough_information}issue_typeβ the 12 allowed valuesobject_partβ the list for that specific object type (a laptop cannot have afront_bumper)risk_flagsβ;-separated from the allowed set, ornonesupporting_image_idsβ only IDs that exist in that claim'simage_paths, ornoneseverityβ {none,low,medium,high,unknown}
Stated plainly rather than left to be found.
- n = 20. One labelled row is 5 percentage points. Every accuracy figure here is low-resolution, and the per-class F1 for
contradictedrests on five examples. contradictedrecall of 0.200 is the headline weakness, not a rounding artifact β see the analysis above.risk_flagsat 15% reflects an over-eager escalation cascade, not a model failure.- No offline path. Without a working API key the pipeline cannot produce output at all.
- No automated test suite. Correctness rests on the evaluation harness alone.
gemini-flash-lite-latestis a moving alias, so exact reproduction over time is not guaranteed.
MIT β see LICENSE. The dataset/ corpus is provided by HackerRank for the Orchestrate challenge and remains theirs.
Adarsh Dwivedi
Built in 24 hours for HackerRank Orchestrate, June 2026. See also: Message Notification Router β the August edition.