WorkWinland Foods
Client story
The claims nobody needs to open
Every short-ship deduction at Winland Foods went to a person, though most investigated claims proved real. We built a model that scores each new claim overnight in Snowflake, clears only the sure ones and was graded on months it never saw.
- Client
- Winland Foods
- Industry
- Food manufacturing
- Timeline
- 5 weeks
- Built with
- Snowflake, Snowflake Model Registry, Snowflake Tasks

The problem
A truck leaves a Winland Foods plant loaded with pasta, sauce and salad dressing. By the store, the count no longer matches the invoice. Sometimes a case went missing. Sometimes it came on a later truck or was returned.
Weeks later the retailer pays the invoice minus what it says never arrived. That deduction is a short-ship claim. Every one went to a person, seven on a typical day and 51 on the busiest.
Only about one in four claims ever got a verdict, 1,821 of 7,599. Nearly three in four of those confirmed a real shortage, so much of the queue asked people to confirm what was usually true.
What we built
Winland’s data team works in SQL and Snowflake, so everything runs there, built for that team to own after we leave.
- At 4:00 AM, half an hour after the SAP sync lands, a scheduled task rebuilds every signal from seven SAP tables.
- It checks the data, scores each new claim through the Snowflake Model Registry and writes the scores to one table, all in 53 seconds.
- Each morning Winland’s UiPath bot reads that table. Claims above the bar clear on their own. The rest reach the team, likely invalid first.
- If the night’s data shrinks by more than a fifth or key signals go missing, the run stops and logs why.
- Each score is written once, the night it is made, so every decision can be audited against exactly what the model knew.
- The bar lives in the UiPath bot, outside the model. The business can move it any morning without touching the pipeline.
How we know it works
An earlier vendor had reported a score of 0.95, where 0.5 is a coin flip and 1.0 is perfect. Its test shuffled every month into one pile, so the model could study next month’s outcomes while predicting last month’s claims.
We graded ours the way it would be used. It learned from claims filed before July 2025 and predicted the months after. The honest score was 0.81. The old shuffled test on our data gave 0.94, so the gap came from the exam, not the data.
The data hid traps too. Nearly half the invoice lines were zero-quantity batch rows. An “Accepted” verdict meant an overage, not a shortage. One signal recorded the outcome instead of predicting it, so we cut it.
- It beats the simplest rule, judging each claim by its customer’s track record, which scores 0.70 on the same honest test.
- Retrained from scratch at three points in 2025, it scored 0.83, 0.81 and 0.83 on the months after each.
- The deductions team had reviewed 21 claims on its own. The model matched 16 of their calls and all five it would have cleared.
- Before launch we rescored all 7,709 claims through the production pipeline. The scores matched development with a correlation of 0.998.
Results
- of claims clear automatically at the recommended bar
- 28%
- of cleared claims were real shortages, on months the model never saw
- 92%
- in deductions cleared through the model at the first month-end close
- $174K
- from kickoff to the first production run
- 25 days
What changed
At the recommended bar, the model clears 28% of claims without anyone opening them. On unseen months, 92% of the claims it cleared were real shortages. For its first month-end close, Winland set the bar a full point stricter, trading volume for certainty. $174K in deductions cleared through the model.
It also measured a hunch. The strongest signal is the shipping plant. Among the busiest plants, the share of shortage claims that hold up runs from 25% to 96%.
Where it stops
- Large claims are harder. Above about $2,200 the score drops to 0.74, so we recommended a dollar cap that keeps the largest claims with a person.
- The answer key has flaws. One confirmed shortage had been denied by the warehouse, with no notes behind it. We recommended a flag for verdicts that feel wrong.
- The share of real claims moves between 63% and 78% by quarter, so we recommended retraining on fresh verdicts every quarter.
What we delivered
- A point-in-time feature pipeline in Snowflake SQL, rebuilt nightly from seven SAP tables
- A classifier in the Snowflake Model Registry, tested on unseen months and against the team’s reviews
- Drift monitoring and a view that checks cleared claims against verdicts as they arrive
- A retraining notebook so the data team can refresh the model on fresh verdicts
Technologies
- Snowflake
- Snowflake Model Registry
- Snowflake Tasks
- scikit-learn
- SAP
- UiPath
More work
Client work
The schools that actually graduate therapists
Hand & Stone needed more licensed massage therapists. We built research agents on a deterministic spine that grew its school list to 1,193, sourced every figure and ranked the schools by what matters to its spas.
Read the story
Client work
The model writes the words. Code owns the numbers.
Acceleration Partners’ account managers wrote every client’s weekly report by hand. We built an AI engine to write them for 43 client programs, where code owns every figure and each one was proven against two outside sources.
Read the story
Bring us the project that matters most.
Tell us what you’re building and what’s in the way. You’ll hear back within 24 hours.


