Field Notes: Supply Chain Systems Failure Patterns and Fixes | BASZ Group

Field Notes

Every note here is a problem we have seen more than once: what it looks like on the floor, why it is usually happening, what to do in the first week, and what good looks like when it is fixed. Filter by your situation, the system involved, or your industry, or search for the symptom you are living with.

Stabilize & Recovery
Implement & Go-Live
M&A Integration
Planning
WMS
TMS
ERP
EDI/API/FTP
OMS
YMS
Data & Analytics
Planning Systems
Quality Systems
Automation
Grocery & Food
CPG
Medical & Regulated
Automotive
Aerospace
Industrial
Showing 86 of 86 field notes

Q: We just went live on a new WMS and short picks have jumped. Is this normal, and where do we start?

Why it happens: A short pick almost never means the product isn’t in the building. It means the WMS thinks it’s somewhere it isn’t. After a go-live that gap opens from three directions at once: the conversion loaded on-hand by location without a fresh physical count, putaway is confirming to the wrong location (or not confirming at all), and replenishment triggers were set from the old system’s min/max rather than the new pick face sizes. Pickers arrive at an empty slot, take a short, and move on. Nobody adjusts the location, so the next wave sends someone to the same empty slot.

Do this first: Pull yesterday’s short picks and sort by location, not by SKU. If a handful of locations account for most of the shorts, you have a replenishment or putaway problem, not an inventory problem. Walk those locations with a supervisor and check three things: was the slot ever replenished, did the replenishment task get confirmed, and is the pick face min set high enough to survive a full wave. Fix the top ten locations by hand today, then fix the rule that produced them. Do not run a wall-to-wall count yet. Counting before you stop the leak just gives you a clean number that will be wrong again by Thursday.

What good looks like: Shorts are below 0.5% of lines within two weeks, every short generates a location task (count or replenish) before the next wave, and Ops and Finance agree on what “short” means. Grocery and CPG sites with heavy case-pick and mixed pallets should expect this to take a little longer because pick face sizing changes by day of week.

When to bring in help: If shorts are still climbing after ten days and the team is arguing about whether it’s a data problem or a floor problem, that argument is the problem. Send us the short-pick report by location and the replenishment task log and we’ll tell you which one it is.

Stabilize & RecoveryImplement & Go-LiveWMSGrocery & FoodCPGinventory accuracy drift

Q: Our operators are working around the RF screens instead of using them. How do we get them back on the system without slowing the floor down?

Why it happens: Operators don’t bypass a system because they’re lazy. They bypass it because the screen is slower than the job. A putaway flow that needs six scans where the old one needed two, a confirmation that times out on the far end of the building, a pick screen that won’t accept a partial and forces a supervisor override — each of these costs seconds, and seconds multiplied by a shift become a bypass. Once a few people are working off paper or keying transactions in batches at the end of the shift, the system is no longer describing what happened on the floor. Inventory accuracy, labor reporting, and wave planning all degrade from there.

Do this first: Shadow three operators for one hour each, in different zones, and count keystrokes and dead time per transaction. You’ll usually find two or three screens that account for most of the friction. Fix those before anything else: remove confirmation prompts that don’t protect anything, shorten scan sequences, allow the partial and quantity exceptions that the floor legitimately hits. Then check RF coverage in the zones where people report the most retries; a heat map from IT takes an afternoon and often explains the “slow” zone. Finally, measure override rate and late-keyed transactions by operator and by screen so the fix is visible.

What good looks like: Transactions are confirmed in real time (no end-of-shift batch keying), override rate is under 2% and falling, and supervisors spend their time on exceptions instead of chasing paper. The floor uses the system because the system is the fastest way to do the job.

A note on aerospace and medical sites: Bypasses here aren’t just an accuracy issue; they break lot, serial, and cert traceability. Treat any paper workaround in a regulated flow as a stop-the-line event and fix the screen the same day.

Stabilize & RecoveryImplement & Go-LiveWMSoverride behavior

Q: Labels and printing keep slowing down our shipping and receiving. Why does something this basic cause so much damage?

Why it happens: Printing is the last step before product leaves a station, so any hesitation there stops the whole station. The usual causes are unglamorous: print jobs routed through a central server across a slow link, printer definitions that don’t match the label stock loaded, ZPL templates that reference fields the WMS didn’t populate for that order type, and printers that quietly drop offline with no alert. Compliance labels (GS1-128, UDI, customer-specific) add another layer, because a mismatched field doesn’t stop the print; it produces a label that fails at the customer’s receiving dock and comes back as a chargeback weeks later.

Do this first: Log print failures for one week by station, printer, label type, and error. Then fix the top three. Move print rendering local where the network is the issue, standardize label templates so one order type equals one template, and put a simple heartbeat on every printer so a supervisor knows within a minute when one is down. Keep a validated spare on each dock. For customer compliance labels, build a pre-ship validation that checks required fields exist and match the ASN before the label is allowed to print.

What good looks like: Print time is under two seconds at the station, printer outages are known before an operator reports them, and label-related chargebacks trend to zero because the same data feeds the label, the ASN, and the invoice.

After an acquisition: Two label standards in one building is the fastest way to create compliance failures. Pick one template library, map the acquired customers into it, and retire the old one on a date.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationWMSIndustrialMedical & RegulatedCPGlabel/print failures

Q: Every wave starts strong and then stalls because pick faces run dry. How do we get replenishment ahead of picking instead of chasing it?

Why it happens: Replenishment is being triggered by the pick, not planned ahead of it. When a wave releases and the WMS discovers the pick face is short, it creates a replenishment task in the middle of the wave and the picker waits. That works at low volume and collapses at peak. Underneath it you’ll usually find pick face min/max set once at design and never revisited, no separation between “top-off” replenishment (done before the wave) and “emergency” replenishment (done during it), and replenishment tasks competing with picks for the same forklifts and the same aisles.

Do this first: Run a demand-based replenishment ahead of each wave: when the wave is built, calculate total demand by pick location and release top-off tasks before pick tasks. Give replenishment its own labor and its own task priority so it isn’t starved by picking. Then review min/max on the top 200 pick locations against actual daily demand; most sites find the fast movers are sized for a Tuesday, not a Friday. If your WMS supports dynamic slotting or demand-based min/max, turn it on for the A items only and watch it for a week before widening.

What good looks like: Emergency replenishment is under 5% of replenishment tasks, waves complete inside their cutoff without supervisor intervention, and pick face sizing is reviewed monthly with the data instead of when someone complains.

Grocery, CPG, and medical: Date-code and lot rotation rules make this harder because the “right” replenishment pallet is not always the nearest one. Make sure FEFO/FIFO logic is driving the replenishment task, or you’ll trade short picks for rotation write-offs.

Stabilize & RecoveryImplement & Go-LiveWMSGrocery & FoodCPGMedical & Regulatedshort picks

Q: The happy path works fine, but every exception (damaged case, wrong quantity, mixed pallet, missing label) takes ten minutes and a supervisor. How do we fix that?

Why it happens: Most implementations test the happy path thoroughly and the exception path once, in a conference room, with a business analyst clicking the buttons. On the floor, exceptions are 10–20% of transactions and they happen with a truck at the door. If the exception flow needs a supervisor login, a reason code nobody can find, or a trip to a desktop, the floor will invent a shortcut. The shortcut usually means the exception never gets recorded, which is why your inventory adjustments don’t reconcile and your damage claims are late.

Do this first: List the top ten exceptions by frequency from the last two weeks (supervisors know them even if the system doesn’t). For each one, time the current path from “operator hits the problem” to “operator is back to work.” Anything over 90 seconds gets redesigned so the operator can resolve it on the RF device with a reason code, and the review happens afterward by someone whose job is review. Give receiving an “accept and flag” option so a truck can be unloaded while the discrepancy is worked. Route the flagged items to a daily exception queue with one owner.

What good looks like: Any operator can clear any common exception in under a minute without leaving their zone, every exception has a reason code and an owner, and the daily exception review is the place where root causes get killed rather than where blame gets assigned.

Stabilize & RecoveryImplement & Go-LiveWMSGrocery & FoodCPGexception overload

Q: Our inventory accuracy is fine on outbound but falls apart on returns. What’s going wrong in the returns process?

Why it happens: Returns arrive with less structure than any other inbound: mixed SKUs in one carton, no ASN, damaged or missing labels, and a disposition decision (restock, rework, scrap, return to vendor) that hasn’t been made yet. Most WMS configurations force the receiver to make that decision at the door, so product gets received as “good” to get it off the dock and sorted out later. Later never comes. Now the system shows sellable inventory that is actually in a quarantine cage or a pile, and outbound starts short-picking against it.

Do this first: Create a returns staging location that is not available to promise, and receive everything there first. Disposition happens in a second step, by a small team with the authority to decide, and the move to sellable stock is the transaction that makes inventory available. Set a dwell limit on the staging location (48 hours is common) and report against it daily. Make sure the returns reason and disposition flow back to the order and to Finance so credits, restocking fees, and vendor claims are driven by one record.

What good looks like: Returns never touch available inventory until they’ve been dispositioned, the staging area clears daily, and the returns data is clean enough for merchandising and quality to use. After an acquisition, get both companies’ return policies into one disposition matrix before the first combined peak.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationWMSCPGGrocery & Foodinventory accuracy drift

Q: Picks per hour have slowly declined for months even though volume is flat. Could slotting be the cause?

Why it happens: Slotting is usually done once, at go-live, against a forecast. Then the product mix changes, new items get slotted wherever there’s an open location, and seasonal movers stay in the golden zone long after their season ends. Nobody notices because nothing “breaks.” Travel time creeps up a few seconds per pick, and six months later the same headcount is producing 15% fewer lines. The clue is that labor cost per line rises while errors stay flat: the floor is working correctly, just farther.

Do this first: Pull pick frequency by location for the last 60 days and overlay it on the pick path. You’ll see A items in C locations and dead items in prime slots. Re-slot the top 100 movers first, during a slow shift, and measure travel time before and after with the WMS task timestamps. Then put a monthly slotting review on the calendar with a simple rule: any item whose velocity class changes by two classes gets moved. If you have a slotting module, make sure it’s using real pick data rather than the original design assumptions.

What good looks like: Travel is under 40% of pick time, slotting is a standing monthly task with a report attached, and new items are slotted by expected velocity rather than by whichever location is open.

Medical and multi-site: Lot and expiry constraints can pin an item to a zone. Re-slot within the constraint; don’t override it. After a merger, re-slot the combined catalog once as a project rather than letting it settle by accident.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationWMSMedical & RegulatedCPGKPI decline

Q: Product is received and shows as available, but pickers can’t find it because it hasn’t been put away. How do we close that gap?

Why it happens: The receipt transaction made the inventory available the moment it was scanned at the dock, but the product is still on a pallet in the staging lane, sometimes for hours. Order promising, allocation, and wave planning all see “available” and act on it. Then the pick task sends someone to a location that is still empty. The causes are usually a putaway queue that’s being worked in arrival order instead of demand order, a receiving team measured on receipts per hour with no accountability for putaway, and a staging lane that isn’t a real WMS location so nothing is visible until it moves.

Do this first: Make receiving-staged inventory unavailable to allocation until putaway is confirmed. That one setting stops the phantom picks today. Then make staging a tracked location with dwell reporting so you can see how long product sits. Prioritize putaway by open demand (hot receipts first) and give the putaway team a target: 90% of receipts put away within two hours. In an ERP-to-WMS integration, make sure the ERP is not making inventory available from the receipt message alone.

What good looks like: Nothing is promised that isn’t in a pickable location, staging dwell is visible on the daily dashboard, and cross-dock or hot-receipt flows are explicit paths rather than a supervisor pulling pallets out of a lane.

Automotive and industrial: Sequenced or kitted parts often need to skip putaway entirely. Build that as a deliberate cross-dock flow with its own confirmation, not as an exception.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationWMSIndustrialMedical & RegulatedAutomotiveWIP mismatch

Q: Cycle counting has become a full-time job and accuracy still isn’t improving. What should the program actually look like?

Why it happens: Counting is being used as a repair tool instead of a measurement tool. If accuracy is poor, the team counts more, finds more errors, adjusts more, and the errors come back because the process that created them is untouched. Meanwhile the count program is usually built on ABC by value, so the team spends its time on high-value slow movers and never counts the fast-moving pick faces where most of the adjustments actually originate.

Do this first: Change what you count. Count by exception first: every short pick, every empty-location report, every negative balance triggers a count of that location the same day. Then count the fast-moving pick faces on a short cycle (weekly for A items) and the reserve locations on a long one. Track adjustments by reason and by process (receiving, putaway, pick, replenishment, returns) and put the top reason on the daily stand-up. Get Finance to agree on a tolerance so small variances don’t require a full investigation.

What good looks like: Count labor drops as accuracy rises because you are counting where errors happen instead of everywhere, location accuracy is above 98% on pick faces, and the adjustment report tells you which upstream process to fix next. Most sites can run the program with a fraction of the labor once exception-driven counting is in place.

Implement & Go-LiveStabilize & RecoveryWMSGrocery & FoodCPGIndustrialinventory accuracy drift

Q: Our contracted carriers are rejecting more tenders on certain lanes. How do we find out why and fix it before we’re living in the spot market?

Why it happens: Carriers reject tenders when the freight has become harder to run than it was when they priced it. The rate didn’t change, but something else did: tender lead time shrank, pickup windows got tighter, dwell at your dock grew, appointment scheduling started failing, or the lane volume doubled and the carrier’s capacity didn’t. The TMS routing guide keeps offering the freight to the same carrier at the same rate, the carrier keeps declining, and the tender cascades down to a backup at a higher cost. If nobody is watching acceptance by lane, this shows up months later as a freight budget miss.

Do this first: Report tender acceptance by carrier, by lane, by week, and look for the point where it fell. Then look at what changed that week: tender lead time, order cutoff, dock dwell, or volume. Call the carrier’s operations contact (not the sales rep) and ask what makes the lane hard. Fix your side first: tender earlier, honor the pickup windows, get trailers loaded on time. If the lane volume has genuinely changed, re-bid the lane rather than waiting for the annual RFP.

What good looks like: Primary carrier acceptance is above 85% on core lanes, routing guide compliance is reported weekly, and the transportation team knows about a declining lane before it hits the spot market.

Stabilize & RecoveryImplement & Go-LiveTMSIndustrialCPGtender rejections

Q: We implemented load optimization and our freight spend went up. How is that possible?

Why it happens: The optimizer is doing exactly what it was told. If it was configured to minimize linehaul cost, it will build loads that look cheap on paper and expensive in practice: multi-stop routes that trigger accessorials, consolidation that pushes orders past the customer’s delivery window, and carrier selections that ignore acceptance rates. The other common cause is data: rates loaded without fuel surcharges or accessorials, transit times that don’t reflect reality, and equipment constraints that weren’t modeled, so the “optimal” load gets rejected and re-planned manually at a worse rate.

Do this first: Take last month’s freight invoices and compare planned cost to invoiced cost by load. Categorize the variance: accessorials, spot premium after a rejected tender, re-planned loads, and rate errors. That tells you whether the optimizer’s plan is wrong or whether execution is drifting from the plan. Then check the objective function and constraints against what the business actually wants (total landed cost including service failures, not linehaul). Load the accessorial schedule and real transit times. Most teams find that two or three configuration choices explain the increase.

What good looks like: Planned-to-invoiced variance is under 5%, the optimizer’s plan is executed as built more than 90% of the time, and Finance and Transportation are looking at the same cost per shipment.

Stabilize & RecoveryTMSinvoice variance

Q: We updated the routing guide after a carrier bid and on-time delivery fell. What did we miss?

Why it happens: A routing guide change is a network change. The new primary carrier on a lane may have a longer transit time, a different pickup cutoff, or fewer trucks in that region on Fridays. If the transit table in the TMS wasn’t updated to match the new carrier, promise dates are still being calculated from the old carrier’s performance. Orders ship on time and arrive late. On retail lanes with must-arrive-by dates, that becomes a chargeback.

Do this first: Compare OTIF by lane before and after the change and isolate the lanes that dropped. For each one, check three things: was the transit time updated for the new carrier, does the new carrier’s pickup cutoff match your order cutoff, and is the carrier accepting on the days you tender. Fix the transit table first because that repairs the promise date. Then talk to the carrier about the specific days that are failing. Going forward, make transit-time updates a mandatory step in the routing guide change process, and run a 30-day watch on every changed lane.

What good looks like: Routing guide changes come with a transit and cutoff review, OTIF is monitored by lane for 30 days after any change, and Sales knows before the customer does when a lane is at risk.

Grocery and CPG: Retailer must-arrive-by windows leave no room for a transit-table error. Load retailer-specific delivery windows into the TMS so the promise date is calculated against the window, not just the carrier’s transit.

Implement & Go-LiveStabilize & RecoveryTMSGrocery & FoodCPGKPI decline

Q: Mode and carrier decisions live in one planner’s head. The TMS suggests something and they override it. How do we get the rules into the system?

Why it happens: The planner is usually right, which is the problem. They know that a certain customer won’t accept LTL, that a lane looks cheap by truckload but pickups are always late, that a shipper prefers a specific carrier for a fragile SKU. None of that is in the TMS, so the system’s recommendation looks wrong and the planner overrides it. Over time the overrides become the process, the TMS becomes a rate lookup, and the day the planner is out the operation stalls.

Do this first: Pull the override log for 30 days and sit with the planner for two hours. Every override has a reason; write it down. Most of them fall into a short list: customer constraints, equipment constraints, carrier performance on a lane, and consolidation rules. Encode those as constraints and preferences in the TMS one category at a time, then measure override rate as it falls. Where a rule can’t be encoded, at least require a reason code on the override so the knowledge is captured.

What good looks like: Override rate is under 10% and every override has a reason, a new planner can run the desk in a week using the system’s recommendations, and the tribal rules have become documented business rules that Sales and Customer Service can see.

Implement & Go-LiveStabilize & RecoveryTMSGrocery & FoodCPGIndustrialexpediting

Q: We can’t bill or close claims because PODs are missing or late. How do we get delivery confirmation under control?

Why it happens: POD is usually the last thing anyone integrates. Carriers send it by email, portal, EDI 214, or not at all; drivers capture signatures on a device that doesn’t sync; customers sign a paper the driver keeps. The TMS marks the shipment delivered based on a status message, but Finance won’t invoice without the document, and Customer Service can’t dispute a short-delivery claim without it. Days-sales-outstanding grows and the claims backlog grows with it.

Do this first: Measure POD availability: percent of deliveries with a POD attached within 24 hours, by carrier. Rank the carriers and go after the bottom five, because they usually account for most of the gap. Set a POD requirement in the carrier agreement with a timeline and a consequence. Route inbound PODs (email, portal downloads, EDI 214 with documents) into one queue that attaches them to the shipment automatically, and put an exception report on deliveries past 48 hours without a document. Give Finance a clear rule for when a status message is sufficient to invoice and when a document is required.

What good looks like: POD is attached to more than 95% of shipments within a day, invoicing doesn’t wait on documents, and claims are resolved with evidence rather than negotiated.

Implement & Go-LiveStabilize & RecoveryTMSIndustrialCPGinvoice variance

Q: Consolidating into multi-stop loads saved freight money but our deliveries are late. How do we keep the savings without losing service?

Why it happens: A multi-stop load is only as reliable as its earliest stop. If the first delivery dwells two hours at the dock, every stop after it slides. The optimizer doesn’t know that a particular customer takes ninety minutes to unload, that a stop has a two-hour receiving window, or that the drive time between stops is optimistic. It builds a load that is feasible on paper and infeasible on a Wednesday. The savings are real; the service cost simply isn’t in the model.

Do this first: Pull stop-level arrival and departure data for the last 30 days and calculate actual dwell by customer location. Load those dwell times and the customers’ receiving windows into the TMS as constraints. Then cap stops per load for the lanes that are failing and sequence stops by window, not by distance. Review the loads that ran late and check whether the plan was infeasible from the start; if it was, the constraint data is the fix. If it was feasible and still failed, the carrier is the conversation.

What good looks like: Multi-stop loads deliver on time above 95%, stop dwell is a tracked metric by customer, and the savings survive because the plan reflects how receiving actually works.

Medical and post-merger networks: Temperature-controlled or regulated product limits how long a load can be on the road. Model that as a hard constraint. After an acquisition, don’t merge two customer bases into shared multi-stop loads until you have dwell data on the new customers.

Stabilize & RecoveryM&A IntegrationTMSMedical & RegulatedCPGKPI decline

Q: Detention charges have doubled and nobody can point to a single cause. Where do we look?

Why it happens: Detention is what a carrier charges you for a problem in your building. The driver arrived and waited because the door wasn’t ready, the load wasn’t built, the paperwork wasn’t done, or the appointment was overbooked. A spike usually means something upstream changed: a new wave schedule that finishes loads later, a receiving team that lost headcount, a yard that lost visibility of which trailer is at which door, or a customer who quietly changed their receiving hours. The invoices arrive weeks later, so by the time Finance sees it the cause is hard to trace.

Do this first: Get the detention detail from the carriers by shipment (date, location, arrival, departure) rather than the monthly total. Sort by location and by hour of arrival. If it clusters at your dock, look at dwell by door and by shift and tie it to the yard and dock schedule. If it clusters at customers, load their actual receiving windows and dwell into the TMS and stop tendering into windows they can’t honor. Put a daily dwell report in front of the dock supervisor, and give someone ownership of detention as a number, not just as an invoice.

What good looks like: Detention is known the day it happens, not the month after, dock dwell is under the carrier’s free time on more than 90% of loads, and every detention invoice can be tied to a cause that someone owns.

Stabilize & RecoveryTMSGrocery & FoodCPGMedical & Regulateddetention fees

Q: We went live on a TMS and tenders aren’t reaching carriers, or carriers aren’t responding. What do we check first?

Why it happens: Tendering is a handoff between three systems that were tested separately: your TMS, the EDI or API connection, and the carrier’s system. Week one exposes the things a test plan can’t: the carrier’s SCAC in your routing guide doesn’t match the one on their EDI profile, the 204 tender is going out but the 990 response is landing in a mailbox nobody watches, tender expiry is set to two hours on a lane the carrier only checks twice a day, or the carrier was never told that tenders would now arrive electronically. Each one looks like “the carrier didn’t respond” and the planner starts phoning.

Do this first: Take the ten most important lanes and trace one tender end to end: did it leave the TMS, did the carrier receive it, did they respond, and did the response come back and update the load. Do this with the carrier on the phone. Fix the connectivity and identifier problems first; they are the fastest to close. Then set tender expiry times by carrier based on how they actually work, and confirm that every carrier in the routing guide has an active, tested connection. For the first two weeks, run a daily tender status report and call anything unanswered after half the expiry window.

What good looks like: Tender-to-response is measured by carrier, the routing guide only contains carriers with a working connection, and planners are handling exceptions rather than re-keying tenders.

Stabilize & RecoveryImplement & Go-LiveTMSIndustrialhandoff failures

Q: Someone changes a UOM, a vendor code, or a routing in the ERP and the warehouse or plant stops. How do we prevent that?

Why it happens: Master data in the ERP is upstream of everything: the WMS, the TMS, the labeling system, the planning engine, and every integration between them. A change that looks harmless in the ERP screen (a new pack size, a changed ship-from, an item made inactive) propagates through interfaces that expect the old value. The WMS rejects the order, the label fails validation, the planning run drops the item, and each downstream team sees a different symptom. Nobody connects it to the master data change because the person who made it had no reason to think it mattered.

Do this first: Identify which master data fields are consumed by which downstream systems, even if it’s a spreadsheet. Then put a change control on the fields that matter: item UOM and pack, vendor and customer identifiers, ship-to and ship-from, item status, and routing or BOM changes. Change control doesn’t mean a committee; it means the change is logged, the downstream owners are notified, and it takes effect on a date rather than immediately. Add a nightly integrity report that catches the common failures (items with no WMS record, orders referencing inactive items, UOM conversions that don’t multiply out).

What good looks like: Master data changes are visible before they take effect, the downstream teams know what changed when something breaks, and the integrity report is empty most mornings.

After a merger: Two ERPs with two item masters are the single largest source of execution failures in the first year. Build the cross-reference and the governance before you connect the systems, not after.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationERPIndustrialGrocery & FoodAerospacecascading outages

Q: The ERP promises dates that the warehouse and carriers can’t hit. Sales trusts the date, the customer plans around it, and we expedite to save it. How do we fix the promise?

Why it happens: Available-to-promise in the ERP is usually calculated from inventory and a fixed lead time. It doesn’t know that the DC’s order cutoff is 2 p.m., that the carrier on that lane doesn’t pick up on Fridays, that the item is in receiving-staged inventory and not yet pickable, or that the site is already at capacity for the day. So the promise is technically correct and operationally impossible. The warehouse finds out when the order drops, and the only way to hit the date is expedited freight.

Do this first: List the constraints that actually govern when an order can ship: site cutoffs, pick and pack capacity, carrier pickup schedules, transit times by lane, and inventory status rules. Then make the promise calculation honor them, either in the ERP’s ATP configuration or by having the OMS or TMS supply the ship and delivery dates back to the ERP. Start with cutoffs and transit times because they fix the most dates for the least effort. Track promised-versus-actual by reason so you can see which constraint is breaking most often.

What good looks like: The promise date is one the operation can hit without expediting more than 95% of the time, expedite spend is reported against the reason it was needed, and Sales stops adding their own buffer because they trust the system.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationERPexpediting

Q: We have inventory and still generate backorders. The allocation logic seems to be the problem. What should we look at?

Why it happens: Allocation rules decide which orders get inventory when there isn’t enough for everyone. When they’re wrong, they reserve stock for orders that won’t ship for a week while an order due today goes on backorder. The common culprits: allocation at order entry (first-come, first-served) instead of at release, reservations that never expire when orders are changed or cancelled, safety stock that’s protected from allocation even when the order is for the customer it was meant to protect, and inventory in a status the ERP treats as unavailable (quality hold, returns, receiving-staged) that is actually fine to ship.

Do this first: Pull the backorder report and, for each backordered line, check whether the item had available inventory anywhere in the network at the time. If it did, find out why it wasn’t allocated: reserved to another order, wrong status, wrong site. That analysis usually points to two or three rules. Move allocation closer to release, expire stale reservations nightly, and make the status rules explicit. If your priority logic favors order date over customer need date, change it.

What good looks like: Backorders exist only when the network is genuinely out of stock, reservation age is monitored, and Customer Service can explain any backorder in a sentence.

Medical and CPG: Lot-specific and customer-specific reservations (contract inventory, retailer-dedicated stock) need to be modeled explicitly or they’ll either starve everyone else or get consumed by the wrong order.

Stabilize & RecoveryImplement & Go-LiveM&A IntegrationERPCPGMedical & Regulatedexcess & stockouts

Q: Orders come in cases, the warehouse picks eaches, the customer receives pallets, and the invoice is wrong. How do we get unit of measure under control?

Why it happens: Every system in the chain has its own idea of the unit: the customer orders in their UOM, the ERP converts it, the WMS picks in its own, the label prints another, and the invoice uses a fourth. As long as the conversion factors agree, it works. They drift when a supplier changes a pack size, a new item is set up without a full conversion table, or a customer’s EDI order uses a UOM code your map treats as something else. The result is a quantity that is right in one system and wrong in the next, and an exception that lands on whoever notices first.

Do this first: Run an item master audit: every active item needs a base UOM and a complete conversion hierarchy (each, inner, case, layer, pallet) with dimensions and weights. Items with gaps are the ones generating exceptions; fix them first. Then agree, across Ops, Sales, and Finance, which UOM is the system of record for inventory and which for pricing. Put a validation on inbound orders and inbound receipts that rejects a UOM the item doesn’t support, rather than letting it convert to something wrong. Finally, add pack-size change to your master data change control so a supplier change updates every system on the same day.

What good looks like: One conversion table drives the order, the pick, the label, the ASN, and the invoice, UOM-related exceptions are near zero, and a pack change is a scheduled update rather than a surprise.

Implement & Go-LiveStabilize & RecoveryERPCPGIndustrialMedical & Regulatedexception overload

Q: Since go-live, Finance can’t reconcile inventory and WIP between the ERP and the floor. Month-end takes weeks. What’s broken?

Why it happens: Before go-live, the reconciliation problems were small enough to absorb. Go-live changes the volume and the timing. Transactions post in a different order, interfaces fail and get replayed, adjustments are made in the WMS without a reason that Finance can book, WIP is consumed at a different point than the ERP expects, and goods-receipt timing no longer matches invoice timing. Each variance is small, but there are thousands of them and nobody owns the bridge between the operational system and the ledger.

Do this first: Stop trying to reconcile the month. Reconcile one day. Pick a single day’s transactions and trace every inventory movement from the WMS or MES to the ERP posting: what posted, what didn’t, what posted twice, what posted to the wrong account. That exercise exposes the handful of transaction types that are causing the variance. Fix those interfaces and reason codes, then run a daily reconciliation with tolerance so month-end is a review of small known items rather than a discovery.

What good looks like: Inventory and WIP reconcile daily within an agreed tolerance, every WMS adjustment has a reason code that maps to a ledger account, and month-end close returns to its pre-go-live duration.

After a merger: Two charts of accounts and two costing methods make this harder. Map the operational transactions to the target ledger before cutover and reconcile both entities separately until the mapping is proven.

Implement & Go-LiveM&A IntegrationERPIndustrialMedical & RegulatedGrocery & FoodWIP mismatch

Q: We find out an order or ASN never arrived only when the customer calls. How do we make EDI failures visible before they cost us?

Why it happens: EDI is a send-and-hope protocol unless you close the loop. You send an 856 ASN or an 810 invoice, the VAN or AS2 connection accepts it, and everyone assumes it arrived. The trading partner’s system rejected it on a validation error and sent a 997 functional acknowledgement saying so. That 997 landed in a folder nobody monitors, or was never requested, or came back three days later. Meanwhile the customer’s receiving dock has no ASN, the invoice is aging in nobody’s AP queue, and the first sign is a chargeback or a missed payment.

Do this first: Turn on 997 tracking for every outbound document type and every partner, and build one report: documents sent, acknowledged, accepted, rejected, and unacknowledged past the partner’s SLA (usually 24 hours). Assign one person to work that report every morning. For inbound documents, do the same in reverse: what did you receive, what did you acknowledge, and what did you reject and why. Then go partner by partner on the rejections; most are a handful of repeating validation errors (missing pack level, bad GS1 label, wrong PO reference) that a mapping fix eliminates.

What good looks like: Every outbound document has a known status within 24 hours, unacknowledged and rejected documents are worked daily by name, and the customer never tells you about a missing ASN because you already re-sent it.

Retail and CPG: Retailer ASN chargebacks are almost entirely preventable with acknowledgement tracking. The cost of the monitoring is trivial compared with one month of compliance deductions.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPCPGIndustrialcascading outages

Q: Every customer wants something slightly different on the ASN and the label. We keep failing compliance audits. How do we manage this without a map per customer?

Why it happens: Retailer and distributor compliance guides are long, change without much notice, and are enforced at the receiving dock by someone who doesn’t care about your mapping. A carton label that is fine for one customer fails another because the SSCC placement is different or a field is missing. The ASN pack structure that works for a full pallet fails for a mixed one. Because compliance data comes from four systems (ERP for the order, WMS for the pack, labeling for the barcode, EDI for the document), a change in one drifts away from the others and nobody sees the mismatch until the deduction.

Do this first: Build a compliance matrix: one row per customer, one column per requirement (ASN hierarchy, label format, timing, PO reference rules, carton and pallet ID rules). Most of the columns are shared; the differences are a short list. Then make the WMS pack data the single source for both the label and the ASN, so they can’t disagree. Add a pre-ship validation that checks the specific customer’s rules before the shipment is closed. Work the chargebacks report monthly by customer and by reason, and fix the top reason each month.

What good looks like: Compliance deductions are under 0.1% of sales to the customer, a new customer requirement is a matrix update rather than a project, and the label, the ASN, and the invoice are generated from the same shipment record.

Medical distribution: Add lot, expiry, serial, and UDI to the matrix. A compliance failure here is a traceability failure, which is a much bigger conversation than a deduction.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPGrocery & FoodCPGMedical & Regulatedchargebacks

Q: A trading partner changed something in their EDI and our orders started failing or, worse, processing wrong. How do we protect ourselves?

Why it happens: Partners change maps. They add a segment, move a reference number, start sending a new qualifier, or change how they express a date. If your map is strict, the document rejects and you notice. If your map is lenient, the document processes with the wrong value in the wrong field, and you don’t notice until a customer receives the wrong quantity or an invoice references the wrong PO. The second case is worse and more common, because most maps are built to be forgiving so that go-live isn’t delayed by partner differences.

Do this first: Add business validation, not just syntax validation, on inbound documents: does the PO exist, does the item exist for this customer, is the quantity within a plausible range, does the ship-to match the customer’s known locations. Route failures to an exception queue with the raw document attached. Then build a simple change-detection check: a daily comparison of segment and qualifier usage by partner against the previous week, which flags a new pattern before it does damage. Finally, get on the partner’s change notification list, and make your own partner-facing changes announced with the same courtesy.

What good looks like: A partner change surfaces as an exception with a clear reason, not as a wrong shipment, and the EDI team can tell you what changed and when for any partner in the last 90 days.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPGrocery & FoodIndustrialdata trust loss

Q: ASNs arrive after the truck, or with the wrong contents, and receiving grinds to a halt. How do we get ASN-driven receiving to work?

Why it happens: ASN-driven receiving assumes the ASN arrives before the trailer and matches what’s on it. When the supplier generates the 856 from the order instead of from the shipment, the contents are what they intended to ship, not what they loaded. When it’s generated at ship confirm but batched overnight, it arrives after a next-day delivery. When the pallet IDs on the ASN don’t match the labels on the pallets, the receiver scans a license plate the system doesn’t recognize and falls back to a blind receipt. Each failure turns a two-minute receipt into a thirty-minute reconciliation, and the truck waits at the door.

Do this first: Measure ASN performance by supplier: percent received before arrival, percent matching the physical receipt, and percent with scannable pallet IDs. Publish it to the suppliers with a target. Set your inbound process to handle the two cases explicitly: ASN present and matched, receive by pallet scan; ASN absent or mismatched, receive to a discrepancy path that lets the truck unload and creates a supplier exception. For your own outbound ASNs, generate them from the confirmed shipment and send them within minutes of the truck leaving, so your customers don’t have the same problem with you.

What good looks like: Over 90% of inbound is received against a matching ASN in a single scan per pallet, supplier ASN compliance is scored and reviewed, and receiving dwell is decoupled from paperwork.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPCPGGrocery & Fooddetention fees

Q: Our EDI passes every validation test and the business results are still wrong: wrong quantities, wrong dates, wrong ship-tos. What did testing miss?

Why it happens: EDI testing certifies that the document is well formed and that the fields land where the spec says they should. It doesn’t certify that the business meaning survived. A date qualifier mapped to “requested ship” when the partner meant “requested delivery” shifts every order by the transit time. A quantity mapped in eaches when the partner sends cases multiplies the order. A ship-to that defaults to the customer’s headquarters when the store number isn’t recognized sends product to the wrong building. Every one of these passes the test harness and fails on the first real order.

Do this first: Take the last 50 orders from each major partner and reconcile the business fields, not the syntax: order quantity versus shipped quantity, requested date versus the date the operation used, ship-to on the document versus where it was delivered. Any systematic difference is a mapping semantics error. Fix it in the map, then add a business-rule validation so the same class of error is caught on entry. Going forward, test every new partner with a set of real historical orders, with someone from Ops reviewing the outcome in the WMS, not just the EDI team reviewing the document.

What good looks like: New partner testing includes an operational review of what the orders would do in the warehouse and on the invoice, and semantic mapping errors are caught by validation rather than by the customer.

Implement & Go-LiveEDI/API/FTPexcess & stockouts

Q: Our system-to-system integrations work until they don’t, and nobody knows they’ve stopped until the floor complains. How do we see what the integrations are doing?

Why it happens: Integrations are built and tested for function, then deployed with logging that was designed for developers. There is no view of message volume by interface, no alert when volume drops to zero, no queue depth, no age of the oldest unprocessed message, and no way to trace one order from ERP to WMS to carrier. When something breaks, IT looks at logs, Ops looks at the floor, and both are right about what they see. The integration has been silently failing for hours.

Do this first: For each integration, define four numbers and put them on one screen: messages in, messages out, errors, and age of the oldest unprocessed message. Then set two alerts: volume drops below the expected baseline for the time of day (a silent failure), and oldest-message age exceeds a threshold (a backlog). Give every message a correlation ID that ties it to the business document (order number, shipment number) so support can trace one order across systems in minutes. This is a few days of work and it changes the conversation from “is the integration down?” to “interface X is 40 minutes behind and here is why.”

What good looks like: Ops and IT look at the same integration dashboard, silent failures are detected in minutes rather than hours, and the first question on any floor issue is answered by a lookup rather than a meeting.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPIndustrialbacklog growth

Q: We’ve had duplicate orders and even duplicate shipments from integration retries. How do we make that impossible?

Why it happens: When a message is sent and no response comes back within the timeout, the sender retries. That is correct behavior. If the receiver processed the first message and only the response was lost, the retry creates a second order, a second pick, a second shipment. The fix is idempotency: the receiver recognizes that it has already processed this message and returns the original result instead of doing the work again. Many integrations don’t have it because the happy path never needed it, and the first time it matters is at peak, when timeouts are most common.

Do this first: Give every message a unique idempotency key (the business document number plus a version works for most cases). On the receiving side, store processed keys and reject or return-cached-result for any repeat. Do this for order creation, shipment confirmation, and inventory adjustments first because they cause the most damage when duplicated. Then audit the last 90 days for duplicates you didn’t catch: same order number or same customer PO with two internal orders, same shipment with two carrier tenders. Clean them up and reconcile with Finance.

What good looks like: Retries are safe by design, duplicates are rejected at the boundary with a logged reason, and the integration team can turn retry policy up during peak without fear.

Stabilize & RecoveryImplement & Go-LiveEDI/API/FTPbacklog growth

Q: One slow system causes every integration behind it to back up, and recovery takes hours. How do we stop a single timeout from taking down the whole chain?

Why it happens: Synchronous chains fail together. The OMS calls the ERP, which calls the WMS, which calls the labeling service. If the labeling service takes ten seconds instead of one, every call in front of it waits, threads are exhausted, and the OMS starts timing out on requests that had nothing to do with labeling. When the slow component recovers, all the queued and retried requests arrive at once and knock it over again. The floor experiences this as “the system is down” for an hour, and then “everything is slow” for the rest of the shift.

Do this first: Identify which calls need to be synchronous (the operator is waiting for an answer) and which can be asynchronous (the answer can arrive later). Move the asynchronous ones to a queue so a slow downstream doesn’t block the upstream. For the synchronous ones, set short timeouts, add a circuit breaker so a failing service is bypassed rather than hammered, and give the operator a defined fallback (queue the transaction, print from a cache, proceed with a flag). Then implement backoff on retries so recovery doesn’t re-trigger the failure. Test all of this under load, not one message at a time.

What good looks like: A slow system degrades one function, not the whole operation, recovery is automatic and doesn’t create a second outage, and the floor has a known procedure for the ten minutes a component is unavailable.

Regulated and sequenced environments: In medical distribution and automotive sequencing, the fallback path needs the same traceability as the primary path. Design the offline procedure to capture the same data, just later.

Implement & Go-LiveStabilize & RecoveryEDI/API/FTPMedical & RegulatedGrocery & FoodAutomotivecascading outages

Q: Our integrations to carriers, marketplaces, and partners keep getting throttled or blocked. Others flood our systems. How do we manage rate limits on both sides?

Why it happens: Every external API has a rate limit, published or not. If your integration sends ten thousand tracking requests at 6 a.m. because a batch job woke up, the partner throttles you (HTTP 429) or blocks you, and your integration treats that as an error and retries, which makes it worse. On your side, if you expose an API without limits, one misbehaving partner integration can consume all your capacity and slow down your own warehouse. Neither case shows up in testing because test volumes never approach the limit.

Do this first: Inventory your outbound integrations and record each partner’s limit and how they signal it. Add a rate limiter on your side so you never exceed it, spread batch jobs across time, and handle 429 with a backoff rather than an immediate retry. For inbound, put limits on your API by partner and return a clear throttling response so a well-behaved partner slows down instead of failing. Then monitor request volume by partner and by hour, so a change in a partner’s behavior is visible before it becomes an incident.

What good looks like: Throttling responses are rare and handled gracefully, no single partner can degrade your platform, and batch jobs are scheduled around the limits rather than colliding with them.

Implement & Go-LiveEDI/API/FTPMedical & RegulatedCPGexception overload

Q: Our integration error queue has thousands of messages in it. Nobody works it. Is that a problem, and how do we get it under control?

Why it happens: An error queue is only useful if someone owns it. When the same validation error repeats a hundred times a day and nobody has authority to fix the source, the queue fills with noise, the real errors get buried, and eventually the team stops looking. That’s when it becomes dangerous: a rejected inventory adjustment or a failed shipment confirmation sits in the queue for weeks, and the systems drift apart in ways that surface as phantom inventory and unbilled shipments.

Do this first: Classify the queue by error type and count. Usually five error types explain 90% of the volume. For each, decide: fix the source (a mapping or master data problem), auto-resolve (a safe rule), or route to a named owner with a service level. Purge or archive the messages that are older than the business can act on, after confirming with Finance and Ops that the underlying transactions were resolved another way. Then set a rule that the queue is worked to zero daily, with a report of what was resolved and how. Give the queue an owner in Operations, not just IT, because most errors are business errors.

What good looks like: The error queue is empty at end of day, error volume by type is trending down because sources are being fixed, and a message that fails is a signal rather than noise.

Medical: In regulated distribution, a failed transaction that is never resolved is a traceability gap. Treat queue age as a compliance metric.

Implement & Go-LiveStabilize & RecoveryEDI/API/FTPMedical & Regulatedbacklog growth

Q: Our retry logic was supposed to make integrations more reliable, and instead it’s creating duplicates and exceptions. What went wrong?

Why it happens: Retries are added to fix a reliability problem and, without three companion controls, create a correctness problem. The three controls are idempotency (the receiver can safely process the same message twice), backoff (retries are spaced so they don’t compound a load problem), and a retry limit with a dead-letter path (a message that fails repeatedly goes to a human rather than looping forever). Without idempotency you get duplicates. Without backoff you get retry storms. Without a limit you get a message that has been retried ten thousand times and is now blocking the queue.

Do this first: Audit every integration that retries and confirm all three controls exist. Add idempotency keys where they’re missing, starting with anything that creates orders, shipments, or inventory movements. Set exponential backoff with jitter and a sensible cap. Route messages that exhaust their retries to a dead-letter queue with an alert and an owner. Then review the last month of exceptions for duplicates the retry logic created, and reconcile them. Most teams find that a single missing idempotency check on one interface explains most of the damage.

What good looks like: Retries are invisible to the business because they never create a second transaction, retry storms don’t happen, and dead-lettered messages are worked the same day.

Implement & Go-LiveStabilize & RecoveryEDI/API/FTPCPGexception overload

Q: Our overnight file transfers fail sometimes and we don’t find out until the morning shift can’t start. How do we make batch integrations safe?

Why it happens: Batch integrations were designed for a world where nothing happened overnight. A file didn’t arrive, a job failed halfway, or a file arrived empty, and nobody was watching. The morning shift starts with no orders, no ASNs, or yesterday’s inventory. Recovery means re-running the job, which may re-process the half that succeeded, and the yard fills up with trucks waiting for a receiving plan that doesn’t exist yet.

Do this first: Instrument the batch the same way you would an API: expected files by time window, files received, file size and record count versus baseline, and job completion status. Alert on absence, not just on failure; a file that never arrives doesn’t throw an error. Make every job restartable from the point of failure rather than from the beginning, and give the on-call person a runbook they can execute at 3 a.m. Then, for the flows that matter most (orders, ASNs, inventory sync), ask whether they still need to be batch at all. Moving the critical ones to near-real-time removes the overnight risk entirely.

What good looks like: A missing or short file triggers an alert within minutes of its window, jobs are restartable and idempotent, and the morning shift has never been surprised by the previous night.

Implement & Go-LiveStabilize & RecoveryEDI/API/FTPCPGyard congestion

Q: Files transfer successfully but the record counts or totals don’t match, and we only find out when inventory or invoices are wrong. How do we validate what we receive?

Why it happens: A successful file transfer proves that bytes moved, not that the content is complete or correct. Files get truncated mid-write, a source job exports before its own upstream finishes, an encoding change drops a line, or a partner sends the same file twice with different content. Without control totals (record count, sum of quantities, a hash) the receiving system loads whatever it got, and the difference shows up later as an inventory variance or an invoice that doesn’t match the PO.

Do this first: Agree a control record with each partner: a trailer or companion file with record count and one or two sum fields. Validate it before loading anything, and reject the file with a clear reason if it doesn’t match. Where a partner won’t provide controls, compute your own baseline (records per file by day of week) and flag files outside the range for review. Then add a post-load reconciliation: what the file said versus what the system now shows. Keep it simple; a daily email with three numbers per interface catches most problems.

What good looks like: No file is loaded without validation, partial and duplicate files are caught at the boundary, and the reconciliation report is boring because it matches.

Implement & Go-LiveEDI/API/FTPIndustrialdata trust loss

Q: A file didn’t arrive for three days and nobody noticed. Orders piled up at the partner and we shipped late. How do we catch the file that never shows up?

Why it happens: Monitoring is almost always built around errors. A missing file produces no error; it produces silence. The scheduler ran, the folder was empty, the job completed with zero records, and every dashboard was green. Meanwhile the partner’s orders were sitting in their outbound folder because their credentials expired, their job was disabled, or their IP changed and your firewall blocked it. Three days later the backlog is large enough that someone notices.

Do this first: Define an expected schedule for every inbound file (partner, file type, window, minimum record count) and alert on any window that closes without a qualifying file. Zero-record files count as missing unless the partner has confirmed there was nothing to send. Put the alert in front of a business owner, not just IT, because the business owner knows whether three days of nothing from a major customer is plausible. Then work with the top partners on a heartbeat: a small file or message sent on schedule even when there’s no data, so absence is unambiguous.

What good looks like: Every expected file has a deadline and an owner, a missing file is an alert within the hour, and the partner hears from you before they notice the problem themselves.

Aerospace and long-lead supply chains: Missed exchanges with suppliers and MRO partners can silently push a multi-week lead time even further. Treat scheduled file exchange as part of supplier performance, not just IT.

Implement & Go-LiveStabilize & RecoveryEDI/API/FTPCPGAerospacebacklog growth

Q: Our OMS promises orders against inventory that turns out not to be available, and we cancel or short customers. How do we make the promise honest?

Why it happens: The OMS promises against a view of inventory that is optimistic in several small ways. It counts inventory that is received but not put away, on quality hold, reserved for another channel, or sitting in a store’s backroom that hasn’t been counted in a year. It also counts inventory once for every channel that sees it, so a single unit is promised to a marketplace order and a web order at the same time. Each optimism is a fraction of a percent; together, at peak, they become a cancellation rate that customers notice and marketplaces penalize.

Do this first: Define “available to promise” precisely, in writing, with Ops and Finance in the room: which inventory statuses, which locations, what safety threshold by node, how channels share or ring-fence stock. Then make the OMS enforce that definition and reconcile it daily against the WMS. Add a promise buffer on nodes with poor accuracy until their accuracy improves, and stop promising from any node that can’t report inventory within an hour. Track cancellations and shorts by reason and by node, and fix the node with the worst accuracy first.

What good looks like: Cancellation rate is under 1%, the promise is calculated from one inventory definition everyone agrees on, and a node that can’t keep its inventory honest loses the right to promise until it can.

Implement & Go-LiveStabilize & RecoveryOMSCPGchargebacks

Q: Returns get credited before they arrive, or twice, or never. Finance can’t close the books. What should the returns flow look like?

Why it happens: A return is three events that happen at different times in different systems: the customer’s request (OMS), the physical receipt and disposition (WMS), and the credit (ERP or payments). When the OMS issues the credit at request time, Finance is refunding product that may never come back. When the WMS receipt triggers a second credit because the OMS return wasn’t linked, the customer is refunded twice. When the return arrives without a return authorization, it gets received as stock with no credit at all. Each path is a Finance reconciliation item that needs a human.

Do this first: Map the three events and decide, once, which one triggers the credit and which one triggers the inventory. Most operations settle on: credit on receipt for goods, credit on request only for specific low-value or trusted categories. Then make the return authorization the link across all three systems: no receipt without an RA, no credit without a linked receipt, and an unmatched-receipt queue for the returns that arrive anyway. Give Finance a daily reconciliation of returns requested, received, dispositioned, and credited.

What good looks like: Every credit ties to a receipt or to an explicit policy exception, duplicate credits don’t happen, and the returns reconciliation is a report rather than a project.

Implement & Go-LiveOMSinvoice variance

Q: We have customers who get priority, faster shipping, or special handling, but that knowledge lives with Customer Service. Orders get treated wrong. How do we encode service levels?

Why it happens: Service commitments are made in contracts and sales conversations and then communicated to the operation by email, sticky note, and memory. The OMS treats every order the same, so a priority customer’s order waits behind a low-priority one, ships by the wrong method, or is packed without the required paperwork. When the customer complains, the operation is blamed for a rule it was never given.

Do this first: Get Sales, Customer Service, and Operations to write down the service tiers that actually exist: what each tier promises (cutoff, ship method, packaging, documentation, allocation priority) and which customers are in it. It is usually three or four tiers. Encode them in the OMS as customer attributes that drive order priority, carrier selection, and packing instructions automatically. Then measure adherence by tier, and make the tier list the only place a service commitment can be made.

What good looks like: An order carries its service tier from entry to delivery without anyone remembering anything, tier adherence is a reported number, and a new commitment to a customer is a configuration change rather than a conversation.

Implement & Go-LiveOMShandoff failures

Q: Our OMS keeps splitting orders across nodes or across days, so we pay for two shipments and the customer gets two boxes. How do we control splits?

Why it happens: Splitting is what an OMS does when it is optimizing for fill rate and nothing else. If one node has four of five lines and another has the fifth, the OMS sources from both, ships two parcels, and reports a 100% fill. It doesn’t see the second shipping charge, the second packaging cost, or the customer opening two boxes on different days. The same logic splits across days when part of an order is on backorder and “ship what you have” is the default.

Do this first: Pull the split rate by reason: multi-node sourcing, backorder, carrier cutoff, and packaging. Then set sourcing rules that weigh total cost, not just availability: prefer a single node that can fill the whole order even at a slightly longer transit, hold a partial for up to a defined number of hours when the remainder is inbound, and require an explicit customer or tier rule before a backorder split is allowed. Review the rules against the actual cost per order, including second-parcel and customer contact costs, so the trade-off is visible.

What good looks like: Split rate is a managed number with a target by channel, the sourcing rule is explainable to Finance in cost terms, and a customer receiving two boxes is a choice the business made rather than an accident.

Implement & Go-LiveOMSCPGmissed carrier cutoffs

Q: We can’t tell what’s in the yard, which trailers are loaded, or which are empty, so drivers wait and doors sit idle. Where do we start?

Why it happens: The yard is the one part of the operation that usually has no system of record. Trailers arrive, get parked, get moved by a jockey with a radio, and the only record is a whiteboard or a spreadsheet updated when someone remembers. The WMS knows a receipt is expected but not which trailer holds it. The dock supervisor sends a jockey to look. Meanwhile a live-unload driver waits at the gate, a loaded outbound trailer misses its pickup, and detention accrues on both.

Do this first: Before buying anything, establish a yard check twice a shift: walk the yard, record every trailer by spot, carrier, status (loaded, empty, in progress), and contents if known. Reconcile that against the WMS’s expected receipts and outbound loads. This alone usually cuts the searching. Then put the check-in gate on a simple system that captures trailer, carrier, seal, and contents, and update status at every move. If you already own a YMS, the problem is usually that the gate and the jockeys aren’t using it consistently, not that it lacks features.

What good looks like: Any supervisor can answer “where is trailer X and what’s on it” in ten seconds, yard moves are tasks with confirmation, and the WMS and the yard agree on what is waiting to be received.

Stabilize & RecoveryImplement & Go-LiveYMSIndustrialGrocery & Foodyard congestion

Q: Trailers sit at doors long after they’re loaded or unloaded, new trailers can’t get in, and the yard backs up. How do we keep doors moving?

Why it happens: A door is a resource, and nobody is measuring its utilization. A trailer finishes unloading and stays because the paperwork is being worked, or the jockey is on the other side of the yard, or the outbound trailer that should replace it isn’t staged. On the outbound side, a loaded trailer stays on the door because the carrier’s pickup is hours away and no one wants to move it twice. Each of these is reasonable in isolation and together they turn a twenty-door dock into a twelve-door dock.

Do this first: Measure door dwell: time from door assignment to release, split into active (loading or unloading) and idle. Publish it by shift. Set a rule that a trailer leaves the door within a defined number of minutes of completion, and give the jockey team a queue of moves driven by that rule rather than by radio calls. Stage loaded outbound trailers in the yard, not on the door, unless pickup is imminent. Then look at the dock schedule: if idle dwell clusters at shift change or at lunch, that is a staffing pattern, not a yard problem.

What good looks like: Idle door time is under 15% of total door time, trailer moves are prioritized by the dock plan, and detention on your own dock is near zero because nobody waits for a door that is already free.

Stabilize & RecoveryYMSCPGdetention fees

Q: We have a dock appointment system but carriers show up when they want, and our receiving team accommodates them. How do we make appointments real?

Why it happens: An appointment system only works if both sides believe it. Carriers stop honoring windows when the dock doesn’t either: a driver arrives on time and waits three hours because an early arrival was taken first, so next time they arrive early too. Suppliers book a window and ship on a different day because the appointment was made by a scheduler who doesn’t talk to the shipping desk. Receiving accommodates everyone because turning a truck away feels worse than running late. Within a month the schedule describes nothing.

Do this first: Start by measuring adherence: arrivals within window, early, late, and no-show, by carrier and supplier. Share the numbers. Then enforce one rule at a time: on-time arrivals are worked first, early arrivals wait until their window, late arrivals are worked in the next open slot. Give the gate the authority and the information to apply the rule. Match appointment capacity to actual receiving capacity by hour so the schedule is feasible; an overbooked schedule teaches everyone that it doesn’t matter. Review adherence with the top ten suppliers monthly.

What good looks like: Appointment adherence is above 85%, receiving labor is planned from the appointment schedule, and a supplier who consistently misses windows hears about it from Procurement, not just from the dock.

Implement & Go-LiveStabilize & RecoveryYMSGrocery & FoodCPGyard congestion

Q: Our gate data is wrong often enough that we can’t use it for detention disputes, carrier scorecards, or yard planning. How do we get accurate gate data?

Why it happens: The gate is staffed by people who are measured on getting trucks through, not on data quality. The check-in form asks for fields the guard can’t verify (contents, PO number), trailer numbers are keyed by hand and transposed, and check-out is skipped when the line is long. Then the data is used for something that matters (a detention dispute, a carrier scorecard) and it doesn’t hold up, so the operation goes back to phone calls and memory.

Do this first: Cut the check-in to the fields the gate can capture accurately: trailer number, carrier, seal, arrival time, and appointment reference. Everything else comes from the appointment or the ASN. Use scanning (trailer plates, appointment barcodes) wherever possible. Make check-out as fast as check-in and make it mandatory by design: the gate doesn’t open without it. Then audit the gate log against the yard check daily for two weeks and fix the patterns you find. Once the data holds up, use it, and tell the gate team that it is being used; data that matters gets entered more carefully.

What good looks like: Gate timestamps are accepted as evidence by carriers and by your own Finance team, the yard system’s view matches a physical yard check, and detention disputes are resolved with a report rather than an argument.

Implement & Go-LiveYMShandoff failures

Q: Our yard jockeys work off radio calls and instinct. Moves get missed, trailers get moved twice, and doors wait. How do we manage yard moves like tasks?

Why it happens: Yard moves are the only warehouse tasks that aren’t tasks. Every pick and putaway is a system-directed, confirmed, measured unit of work; every yard move is a conversation. So the jockey does what the loudest supervisor asks, the dock plan changes without the jockey knowing, and a trailer that should have been at door 12 at 2 p.m. is still in row C at 3. Nobody can say how many moves the yard does in a day or how long each takes, which means nobody can staff it or improve it.

Do this first: Turn moves into tasks: a request with a from, a to, a priority, and a due time, confirmed on completion. This can start on a tablet or even a shared board before it lives in a YMS. Feed the task list from the dock schedule (next trailer to a freed door) rather than from radio calls. Measure moves per hour and move completion against due time, and use the numbers to decide how many jockeys each shift really needs. Then connect it to the WMS so a completed unload releases the door and generates the next move automatically.

What good looks like: The jockey team works a prioritized queue, move times are known, doors are refilled within minutes of release, and the yard is staffed by data instead of by complaint.

Stabilize & RecoveryYMSCPGdetention fees

Q: We have dashboards for everything and the operation runs exactly the way it did before we built them. What are we doing wrong?

Why it happens: A dashboard is a report with better fonts unless it is attached to a decision. Most dashboards are built to answer “how are we doing?” which is a question that generates nodding, not action. The numbers are lagging (last week’s OTIF), aggregated (network fill rate), and not owned (nobody’s job changes when the number moves). The people who could act don’t look at them, and the people who look at them can’t act.

Do this first: Take the three decisions that matter most this quarter (where to add labor, which lane to re-bid, which supplier to escalate) and ask what number would change that decision and who makes it. Build one view per decision, with the number, its trend, its threshold, and the owner’s name on it. Then put it in the meeting where the decision gets made and take the other dashboards out of that meeting. Measure the dashboard by the decisions it changed, not by the views it got. Most operations need five or six of these views, not fifty.

What good looks like: Every metric on a management dashboard has an owner and a threshold that triggers an action, the weekly ops review runs off the dashboard rather than off a deck, and someone can point to a decision last month that the data changed.

Stabilize & RecoveryPlanningData & AnalyticsKPI decline

Q: Planners, supervisors, and customer service override the system constantly and we have no idea how often or why. Does it matter?

Why it happens: It matters more than almost anything else you could measure, because override rate is the honest score for how well the system fits the operation. Every override is either a place where the system is wrong (a rule, a parameter, a data gap) or a place where a person is wrong (a habit, a shortcut, a favor). You can’t tell which until you count them. Most systems allow overrides without a reason code, so the knowledge disappears the moment the override is saved.

Do this first: Turn on override logging wherever it exists (allocation, carrier selection, forecast, safety stock, pick sequence, quality disposition) and require a reason code from a short list. Report override rate by function, by user, and by reason weekly. Then review the top reasons with the people doing the overriding, without blame: an override that everyone does the same way is a rule the system should have. Encode it. An override that one person does is a training or a control conversation. Watch the rate fall as the rules improve.

What good looks like: Override rate is a standing metric with a target by function, every override has a reason, and the monthly review turns the top reason into a configuration change. In grocery and other high-velocity environments, this is often the fastest way to find where the system and the floor disagree.

Stabilize & RecoveryPlanningData & AnalyticsGrocery & Foodoverride behavior

Q: Ops says fill rate is 97%, Sales says 91%, Finance has a third number. Every meeting starts with an argument about whose number is right. How do we fix this?

Why it happens: All three numbers are correct. Ops counts lines shipped against lines released to the warehouse. Sales counts lines delivered against lines ordered. Finance counts dollars invoiced against dollars ordered. Each team built its metric from the data it had, for the decision it needed, and nobody wrote the definitions down next to each other. The argument isn’t about accuracy; it’s about the absence of a shared dictionary. Until there is one, every dashboard is a debate.

Do this first: Pick the ten metrics that show up in the executive review and write a one-page definition for each: numerator, denominator, source system, timing, exclusions, and owner. Get Ops, Sales, and Finance to sign the page. Where the teams legitimately need different views (Ops needs a warehouse fill rate, Sales needs a customer fill rate), keep both, name them differently, and show them side by side with the bridge between them. Then rebuild the reports from the signed definitions and retire the ones that don’t match.

What good looks like: Every metric in the management review has a published definition, two teams quoting different numbers are quoting different metrics by name rather than disagreeing about one, and the data team spends its time on analysis instead of reconciliation.

Stabilize & RecoveryPlanningImplement & Go-LiveData & AnalyticsIndustrialCPGdata trust loss

Q: We find out we’re behind when the shift ends and the work isn’t done. How do we see a backlog forming while we can still do something about it?

Why it happens: Most operational reporting measures throughput (what got done) and not queue (what is waiting). Throughput looks fine until the moment it doesn’t, because the queue was growing quietly all day. Orders waiting for release, pallets waiting for putaway, trailers waiting for a door, exceptions waiting for review, integration messages waiting to process: each is a queue, and a growing queue is the earliest warning the operation has. If no one is watching it, the first signal is the missed cutoff.

Do this first: List the queues in your operation, from order entry to the carrier pickup, and for each one find where the count already exists in a system. Put the counts on one screen, refreshed every few minutes, with the age of the oldest item in each queue. Then set a threshold per queue that means “act now” (release labor, open a door, escalate an integration) and name who acts. Trend the queues by hour so you can see the pattern: a queue that always grows at 10 a.m. is a scheduling problem, not a surprise.

What good looks like: Supervisors are looking at queues, not just at completions, the shift lead knows by mid-morning whether the day is at risk, and cutoffs are missed by decision rather than by discovery.

Stabilize & RecoveryPlanningData & AnalyticsGrocery & Foodbacklog growth

Q: We do root cause analysis, we find the cause, and the same problem comes back next month. Why doesn’t root cause work stick?

Why it happens: Root cause analysis produces a document; it doesn’t produce a change. The five-whys are done, the cause is named (“receiving didn’t validate the pack”), and the corrective action is a reminder to the team. The process that let the error through is unchanged, the metric that would show recurrence isn’t tracked, and the next time it happens it is treated as a new event. Root cause work only sticks when it ends in a control: a system validation, a required field, a standard work change with an audit, or a metric with an owner.

Do this first: Take the last ten root cause investigations and check each one: is the corrective action a control or a reminder? Convert the reminders into controls. Then change the template: every investigation must end with a control, a metric that would show recurrence, and a date to check it. Keep a single register of open corrective actions, review it weekly, and close an item only when the recurrence metric has stayed clean for a defined period. Report recurrence rate: how many of this month’s incidents are repeats.

What good looks like: Corrective actions are system or process changes rather than emails, recurrence rate falls month over month, and the investigation register is short because things actually get fixed.

Stabilize & RecoveryImplement & Go-LiveData & AnalyticsCPGexception overload

Q: All our KPIs tell us what already happened. By the time OTIF drops, we’ve already failed. What should we be watching instead?

Why it happens: Lagging indicators are easy to define and easy to defend, which is why they fill executive dashboards. OTIF, fill rate, cost per case, and inventory accuracy are all outcomes; by the time they move, the causes are two weeks old. Leading indicators are messier (queue depth, exception rate, override rate, dwell time, supplier ASN timeliness, tender acceptance) and they require someone to decide what they predict. Most teams never do that work, so they manage by looking in the mirror.

Do this first: For each lagging KPI that matters, ask what you would have seen a week earlier. OTIF drops after tender acceptance drops and dock dwell rises. Fill rate drops after replenishment exceptions and putaway dwell rise. Inventory accuracy drops after override rate and unresolved exceptions rise. Pick two or three leading indicators per outcome, put them on the daily operations review with a threshold, and give them owners. Then validate them over a quarter: did they move before the outcome did? Keep the ones that predict and drop the ones that don’t.

What good looks like: The daily review is about leading indicators and the weekly review is about outcomes, the team can name which leading indicator explains each KPI movement, and the executive dashboard shows both.

PlanningStabilize & RecoveryData & AnalyticsIndustrialKPI decline

Q: We carry plenty of safety stock and still stock out on the items that matter. Why isn’t the buffer working?

Why it happens: Safety stock is calculated to protect against one kind of variability, usually demand variability over a fixed lead time, and the stockouts are being caused by another: supplier lead time variability, inventory accuracy, or the buffer sitting in the wrong node. A safety stock set at the DC does nothing for a store that isn’t replenished on time. A buffer computed from forecast error does nothing when the supplier ships two weeks late. And a buffer that the system thinks is on hand but isn’t (because accuracy is 92%) is a number, not stock.

Do this first: Take the last 30 stockouts on A items and classify the cause: demand spike, late supply, inaccurate on-hand, wrong node, or parameter never updated. That classification usually shows that most stockouts weren’t demand variability at all, which means the safety stock formula was solving the wrong problem. Fix the top cause directly (supplier lead time policy, cycle counting on the pick faces, positioning stock at the node that serves demand), then re-parameterize safety stock to include supply variability where it matters. Review parameters quarterly, by exception, rather than leaving them at go-live values.

What good looks like: Safety stock is set from measured demand and supply variability, by node, with a service target that Sales and Finance agreed to, and stockouts on A items are explainable by cause.

PlanningPlanning SystemsGrocery & FoodCPGexcess & stockouts

Q: Supplier minimum order quantities are forcing us to buy far more than we need on slow items. How do we manage MOQ-driven excess?

Why it happens: An MOQ is a supplier’s cost structure pushed onto your balance sheet. For a fast mover it doesn’t matter; the minimum is a few days of demand. For a slow mover, a 1,000-unit minimum against 20 units a month is four years of stock, and the planning system dutifully orders it because the parameter says so. The excess then hides in aggregate inventory metrics, gets written down years later, and nobody connects the write-off to the MOQ.

Do this first: Calculate months-of-supply at MOQ for every purchased item and sort descending. The top of that list is your excess-generating catalog. For each, choose: negotiate the MOQ (often possible when you show the supplier the data), accept a higher unit price for a smaller quantity, consolidate to fewer variants, source elsewhere, or deliberately accept the excess with Finance’s sign-off. Then put an MOQ-to-demand check in the planning system so a new item with a bad ratio is flagged before the first PO. Report projected excess at PO creation, not at write-off.

What good looks like: No PO is placed at more than an agreed months-of-supply without an explicit decision, MOQ negotiations are driven by data, and the annual write-off stops surprising Finance.

PlanningPlanning SystemsCPGIndustrialexcess & stockouts

Q: We invested in inventory optimization software and the recommendations are wrong. Is it the model or the data?

Why it happens: It’s the data, nearly every time. An optimization model takes on-hand, on-order, and demand history as facts and produces a recommendation that is precisely correct for a warehouse that doesn’t exist. If on-hand is overstated by 5% on the pick faces, the model under-orders and you stock out. If on-order includes POs the supplier will never ship, the model waits for supply that isn’t coming. If demand history includes the stockout periods as zero demand, the forecast is biased low. The model amplifies whatever error is in its inputs and delivers it with a confidence interval.

Do this first: Before touching the model, audit the inputs. Compare system on-hand to a physical count on a sample of A items by location. Review open POs against supplier confirmations and close the ones that are dead. Check whether demand history is being corrected for stockouts and promotions. Fix the worst of those three, then re-run the model and compare recommendations with a planner’s judgment on the top 50 items. Where they disagree, find out why. Keep the model’s recommendations advisory until the input quality is proven, and measure input quality as a metric of its own.

What good looks like: Inventory accuracy on the items the model manages is above 98%, open orders reflect what the supplier will actually deliver, demand history is cleaned, and planners trust the model because it is right on the items they can check.

PlanningStabilize & RecoveryPlanning Systemsexcess & stockouts

Q: We improved forecast accuracy significantly and inventory and service barely moved. What are we missing?

Why it happens: The forecast is an input to a decision, and the decision didn’t change. Better accuracy only helps if the replenishment parameters (safety stock, order quantities, review periods) are recalculated from the new error distribution, if the planners stop adding their own buffer on top of the system’s, and if supply actually arrives when the plan says it will. Often the accuracy improvement is at an aggregate level (national, monthly) while the decisions are made at a level the forecast still gets wrong (SKU, DC, week). And sometimes the forecast improved for the items that were already fine.

Do this first: Check where the accuracy improved: by item class, by node, by horizon. Then check whether the safety stock parameters were re-derived after the improvement; if not, the system is still buffering against the old error. Look at planner overrides on the forecast and on the order quantities; if planners are still adding buffer, they haven’t seen evidence that the new forecast holds. Finally, measure supply reliability: a perfect forecast against a supplier who ships late by a week still produces stockouts and excess.

What good looks like: Forecast accuracy is measured at the decision level, safety stock is recalculated automatically as error changes, planner buffer-adding is measured and falling, and inventory turns move with accuracy because the whole chain is connected.

PlanningPlanning Systemsforecast volatility

Q: Nobody reviews the planning parameters. Lead times, safety stocks, and order quantities were set years ago. How much does that cost us and what should the review look like?

Why it happens: Parameters are set at implementation by someone who has since left, with data that has since changed, and they persist because the system keeps running. A lead time of 14 days on an item the supplier now ships in 5 keeps a week of extra inventory forever. A min of 200 on a discontinued-in-all-but-name item keeps buying it. There is no owner for the parameters as a set, no report that flags stale ones, and no routine that questions them. The cost is invisible because it is spread across thousands of items as a few percent each.

Do this first: Run a parameter audit: for each active item, compare the planning lead time to actual receipt lead time over the last year, the safety stock to the demand and supply variability actually observed, and the order quantity to current demand and MOQ. Flag anything off by more than a tolerance. That report is your first review. Fix the top 100 by inventory value, measure the effect, then put a quarterly parameter review on the calendar with an owner and an exception report so the work is a few hours rather than a project.

What good looks like: Every parameter has an owner and a last-reviewed date, the exception report is worked quarterly, and stale parameters are found by the report rather than by the write-off.

PlanningPlanning SystemsIndustrialexcess & stockouts

Q: Sales wants 99% service on everything, Finance wants inventory down 20%, and Planning is stuck in the middle. How do we set a service policy that both sides accept?

Why it happens: Nobody has shown the trade-off in numbers. Sales asks for 99% because they’ve never been shown what it costs, and Finance asks for a cut because they’ve never been shown what service it removes. Planning ends up with a single target for every item, which is the worst of both: too much inventory on the items that don’t need it and not enough on the ones that do. A service policy is a business decision that has been left to a parameter table.

Do this first: Segment the catalog by what it is for: revenue contribution, margin, customer criticality, and supply risk. Then model, for each segment, the inventory required at several service levels. Put that table in front of Sales and Finance together and let them choose. Most businesses settle on something like 99% on the top segment, 95% on the middle, and make-to-order or accept-the-risk on the tail. Write the policy down as a signed document with the segments and targets, load it into the planning system, and review it annually or when the business changes.

What good looks like: Service targets are differentiated by segment and agreed in writing by Sales and Finance, inventory is positioned according to the policy, and the quarterly review discusses whether the policy is right rather than whether Planning is doing its job.

PlanningPlanning SystemsCPGexcess & stockouts

Q: When a supplier misses, it takes weeks to get back on schedule, and we live on expedites in the meantime. How do we shorten recovery?

Why it happens: Recovery is slow because it starts late and nobody owns it. The miss is discovered when the receipt doesn’t arrive, not when the supplier knew they would miss. Then the buyer asks for a new date, the supplier gives an optimistic one, it slips again, and each slip is treated as a new event. There is no recovery plan (what will ship when, in what quantity, by what mode), no daily check on the plan, and no escalation path when it slips. Expediting fills the gap because it is the only lever that works today.

Do this first: Require a written recovery plan from the supplier within 24 hours of any miss on a constrained part: quantity by date, mode, and the cause of the original miss. Track the plan daily until it is complete, and escalate on the first slip, not the third. Give one person ownership of each open recovery, and report recovery duration by supplier so the pattern is visible. Then move upstream: get the supplier to send a schedule confirmation or a shortage notice before the due date, so the miss is known before the truck is empty.

What good looks like: Recovery duration on constrained parts is under two weeks and falling, expedite spend is tied to a named supplier miss with a cause, and the suppliers who recover slowly are having a different conversation at the next review.

Automotive and post-merger supply bases: Sequenced or line-down risk parts need a recovery plan template that includes premium freight authority and an alternate-source trigger. After an acquisition, combine both companies’ supplier performance data before you assume the inherited supply base behaves like your own.

PlanningM&A IntegrationPlanning SystemsAutomotiveIndustrialexpediting

Q: Our planning lead times were entered at setup and never updated. Orders arrive early, late, and rarely on time. How do we get lead times right?

Why it happens: A lead time in the item master is a single number that is supposed to represent a distribution. It was typed in from a supplier quote or a round-number guess, and the actual receipt lead time has since drifted. Planning treats the number as truth, so orders are placed too early (excess) or too late (stockout), and planners compensate by manually pulling in or pushing out, which hides the problem further. Nobody measures actual lead time because the data lives in receipts and POs, which are in two different reports.

Do this first: Calculate actual lead time from PO date to receipt date for every supplier and item over the last twelve months: the average and the spread. Compare that to the planning parameter. Update the parameter to the measured average where the difference is material, and load the variability into safety stock where the system supports it. Then automate the calculation so it refreshes monthly, and give buyers an exception report of items where measured and planned lead times diverge. Share the measured lead times with suppliers; they are usually surprised too.

What good looks like: Planning lead times are derived from measured performance and refreshed automatically, lead time variability is part of the safety stock calculation, and buyers spend their time on the exceptions rather than on manual pull-ins.

Grocery and post-merger: Seasonal suppliers and short-shelf-life items need lead times by period, not a single annual number. After a merger, measure the acquired supply base from receipts rather than inheriting its item master.

PlanningM&A IntegrationPlanning SystemsGrocery & FoodCPGdata trust loss

Q: We send suppliers a scorecard every quarter. Nothing changes. What makes a scorecard actually work?

Why it happens: A scorecard changes behavior when three things are true: the supplier believes the numbers, the numbers connect to something the supplier cares about, and someone follows up. Most scorecards fail all three. The data is disputed (“your receiving didn’t book it for two days”), the score has no consequence (no business shifts, no fee, no preferred status), and the review is an email. The supplier files it. The buyer files it. Next quarter, the same scorecard.

Do this first: Fix the data first: measure on-time against the confirmed date, in-full against the confirmed quantity, and share the underlying line-level data with the supplier before the score so disputes happen before the review, not during it. Reduce the scorecard to the four or five measures that drive your cost (OTIF, lead time adherence, ASN compliance, quality, responsiveness). Then attach consequences: volume allocation between dual sources, preferred-supplier status, or contractual remedies. Hold a quarterly review with the supplier’s operations lead, not just their salesperson, and leave with actions and dates.

What good looks like: Suppliers can reproduce their own score from the shared data, the top suppliers compete for volume on the scorecard, and the quarterly review ends with commitments that are tracked. After a merger, harmonize the scorecard before the first combined supplier review, so nobody is measured two ways.

PlanningM&A IntegrationPlanning SystemsCPGsupplier noncompliance

Q: Every supplier record looks different: different address formats, missing lead times, inconsistent payment terms and contacts. It’s breaking planning, receiving, and payables. How do we clean it up without a year-long project?

Why it happens: Supplier records are created by whoever needed the supplier first, with the fields that person needed. Procurement fills in terms, Receiving fills in ship-from, AP fills in remit-to, and nobody fills in lead time, MOQ, or the operational contact. When two companies merge, this doubles: two supplier masters with overlapping suppliers under different codes, and every downstream process (planning, ASN matching, three-way match) has to guess which record is right.

Do this first: Define the minimum viable supplier record: the ten or so fields that planning, receiving, quality, and payables each need to function, with an owner per field. Then run a gap report against the active supplier base (suppliers with POs in the last year) and fix those first; the inactive ones can wait or be closed. Set a creation standard so no new supplier is activated without the minimum record. For a merger, build the cross-reference between the two masters on the active overlap and migrate to one record per supplier on a date.

What good looks like: Every active supplier has a complete record with a named owner per field, supplier creation is a governed process, and planning, receiving, and payables are reading the same record.

Grocery and traceability: Supplier records for food and regulated product need to carry facility identifiers, certifications, and lot-tracking capability. A missing field here is a recall problem, not a data quality problem.

M&A IntegrationPlanningPlanning SystemsGrocery & Foodtraceability gaps

Q: We dual-sourced our critical parts and when the primary failed, the secondary couldn’t deliver either. What did we get wrong?

Why it happens: Dual sourcing reduces risk only if the second source can actually supply when the first can’t. A secondary that has been qualified on paper but has never shipped, that shares a sub-tier supplier or a region with the primary, that has no tooling or no allocation of capacity, is not a second source; it’s a second name. When the primary fails, the secondary needs weeks to ramp, and by then you’ve spent the premium freight you were trying to avoid.

Do this first: For each dual-sourced critical part, ask three questions: has the secondary shipped production quantities in the last year, how long would it take them to reach the required rate from where they are today, and do they share any upstream dependency with the primary. Anything that fails those questions isn’t risk-mitigated. Fix it by giving the secondary a standing share of volume (20% is common) so they stay live, by mapping sub-tier overlap on the parts that matter most, and by writing the switch procedure: who authorizes it, what the secondary commits to, and what the trigger is.

What good looks like: Every critical part has a secondary that is shipping, a documented time-to-rate, and a known sub-tier map, and the switch has been exercised at least once on a low-risk part so the procedure is real.

PlanningPlanning SystemsIndustrialsupplier noncompliance

Q: Suppliers ship the same part in different pack sizes, on different pallets, with different labels, and the warehouse absorbs the cost. How do we get packaging under control?

Why it happens: Packaging is negotiated last, if at all. The buyer secures price and lead time; the pack configuration is whatever the supplier prefers. So the same item arrives as 12-packs one week and 24-packs the next, on a pallet that doesn’t fit the rack, with a label the WMS can’t scan. Receiving repacks it, relabels it, and puts it away as a one-off. None of that labor is charged back to the supplier or visible to Procurement, so the supplier’s price looks good and the total cost is hidden in the warehouse’s cost per case.

Do this first: Have the warehouse log receiving exceptions caused by packaging for two weeks: repack, relabel, non-standard pallet, missing or unscannable label. Cost it out at a loaded labor rate and by supplier. That number is usually large enough to get Procurement’s attention. Then write a packaging specification per item (pack quantity, pallet pattern, label standard, ASN pallet ID) and put it on the PO. Score suppliers on packaging compliance alongside OTIF, and charge back or negotiate on the ones that don’t comply.

What good looks like: Packaging compliance is a supplier metric, receiving exceptions from packaging are near zero, and the warehouse’s cost per case is no longer subsidizing supplier convenience.

PlanningPlanning SystemsIndustrialCPGKPI decline

Q: Our supplier OTIF report says 96% and our plants are still short of parts every week. Which one is lying?

Why it happens: Neither, but the OTIF measure is answering the wrong question. It measures on-time against the last confirmed date, which the supplier changed twice; in-full against the quantity on the PO, which the buyer reduced to match what the supplier could ship. So the supplier delivered exactly what was rescheduled to, and the plant is short against what it originally needed. Averaging across all parts hides the fact that the misses are concentrated on the twenty parts that stop the line.

Do this first: Measure OTIF two ways and show both: against the original need date and quantity, and against the last confirmed one. The gap between them is the reschedule rate, which is the number the plant feels. Then weight the metric by criticality, or simply report it separately for constrained parts. Look at the reschedule history on those parts and find who is changing the dates and why. Often the buyer is absorbing the miss to keep the score green, which is a measurement design problem, not a buyer problem.

What good looks like: OTIF against original need is the primary measure, reschedules are counted and visible, constrained parts have their own report, and the plant’s shortage list and the supplier scorecard finally agree.

PlanningPlanning SystemsCPGAutomotivebacklog growth

Q: Demand at the customer is fairly stable but our orders to suppliers and our production schedule swing wildly. Why does the volatility grow as it moves upstream?

Why it happens: Every stage in the chain reacts to the stage below it with a lag and a buffer. A small uptick at retail becomes a larger order from the DC, which becomes a larger production run, which becomes a larger raw material order, and then the correction runs the same way in reverse. The amplifiers are familiar: order batching, promotions that are planned in Sales but not in Supply, safety stock rules that recalculate every week, allocation during shortages that teaches customers to over-order, and planners at each stage adding their own judgment on top of the signal.

Do this first: Measure it. Plot the coefficient of variation of demand at each stage (customer orders, DC replenishment, production, supplier POs) for a handful of products. The stage where variability jumps is where the amplifier lives. Then remove the amplifier: share point-of-sale or end-customer demand with upstream planning instead of orders, freeze safety stock parameters between reviews, plan promotions in the S&OP cycle, and cap order-size changes week to week unless a planner justifies them. Take one product family through this before scaling.

What good looks like: Variability upstream is close to variability at the customer, planners can see end demand rather than just the next stage’s orders, and the S&OP process changes a supply plan because demand changed rather than because an order did.

PlanningPlanning SystemsCPGGrocery & Foodforecast volatility

Q: Plant shutdowns, DC moves, peak season, a new customer launch: each one seems to catch the supply chain off guard even though everyone knew it was coming. How do we plan for capacity events?

Why it happens: The event is known but not modeled. The shutdown is on the plant calendar and not in the supply plan, so the system keeps assuming normal output. Peak is on the sales calendar and not in the labor plan, so the DC discovers in week one that it needs 40% more people. The new customer is in the CRM and not in the demand plan, so their first order lands as a surprise. Each function planned its part; nobody planned the intersection.

Do this first: Build a twelve-month capacity event calendar as a standing S&OP input: plant maintenance, DC changes, IT cutovers, peak periods, launches, known customer changes. For each event, name the owner and the functions it touches. Then, in the monthly S&OP, review the next 90 days of events with the pre-build, labor, inventory, and transportation decisions each one requires, and make those decisions on a date. Put the events into the planning system as capacity constraints so the plan reflects them automatically.

What good looks like: No capacity event surprises the operation, pre-builds and staffing plans are decided months out, and the post-event review compares the plan to what happened so the next one is better.

PlanningPlanning Systemscascading outages

Q: Our S&OP meeting is a two-hour status update. Everyone presents, nothing gets decided, and the real decisions happen afterward by email. How do we fix the process?

Why it happens: The meeting was designed around presentations instead of decisions. Each function brings a deck, the executive listens, and the decision that needs cross-functional agreement (accept the demand plan and its inventory consequence, choose which customer to short, approve the pre-build) never has a slide with options and a recommendation. So it doesn’t get made. The people who could make it don’t want to decide in a room without the analysis, and the analysis was never asked for.

Do this first: Rewrite the agenda around decisions. Before the meeting, the planning team produces a short list of open decisions, each with options, the trade-off in inventory, service, and cost, and a recommendation. The meeting reviews those, decides, and records the decision and owner. Status is pre-read, not presented. Cut the meeting to an hour. Track decision throughput (decisions made per cycle) and decision latency (how long an issue waited). If a decision can’t be made in the room, it is escalated with a date, not deferred.

What good looks like: Every S&OP cycle ends with recorded decisions and owners, the executive sponsor makes the trade-offs in the room, and functions come prepared to choose rather than to present.

PlanningPlanning Systemsforecast volatility

Q: Sales has a forecast, Finance has a budget, Operations has a plan, and they don’t match. Which number should the supply chain plan to?

Why it happens: Each number was built for a different purpose and never reconciled. The budget was set months ago and is a commitment; the sales forecast is this month’s view and is optimistic; the ops plan is what the plants can make and is conservative. Without a process that brings them together, each function runs its own number, and the gaps become the stockouts and the excess that show up two quarters later. The supply chain plans to the one it trusts most, which is usually the one it built itself.

Do this first: Institute a demand review that produces one unconstrained demand plan (volume by product family by month, with assumptions) and a supply review that produces one constrained supply plan against it. The gap between the two is what the executive S&OP decides about. Then reconcile the demand plan to the budget explicitly, with the gap and its reasons written down, so Finance can see where the risk is instead of learning it later. Keep one planning system of record and stop the spreadsheets from becoming a second one.

What good looks like: There is one demand plan and one supply plan, both owned, both dated, and the budget gap is a managed number with a plan to close it rather than a surprise at quarter end.

PlanningPlanning Systemshandoff failures

Q: Every promotion creates a stockout during the event and a glut afterward. The DC and the carriers are overwhelmed for two weeks. How do we plan promotions properly?

Why it happens: Promotions are planned in Marketing and Sales on a calendar that Supply learns about late, at a volume that is a guess, with a start date that retail partners shift. The supply chain sees a demand spike it wasn’t warned about, expedites to cover it, then holds the inventory that arrived after the event ended. The DC has no labor plan for the surge, the carriers have no capacity commitment, and the post-event overstock sits until it is discounted again.

Do this first: Put promotions into the S&OP demand review with a lift estimate, a confidence level, and a start and end date that Sales owns. Build the pre-build and the labor and carrier plans from that, and set a cutoff after which a promotion can’t be added without a defined expedite budget. After each event, compare planned lift to actual and feed the accuracy back to the next estimate. For the post-event tail, decide the exit plan before the event starts: where the residual inventory goes and who pays for it.

What good looks like: Promotional lift accuracy improves event over event, the DC and carriers see the calendar before the surge, and post-event overstock is a planned quantity rather than a discovery.

Grocery, CPG, and regulated product: Short shelf life and lot-controlled items turn post-event overstock into write-offs. Plan the exit with a date-code view, not just a quantity.

PlanningPlanning SystemsGrocery & FoodCPGMedical & Regulatedexcess & stockouts

Q: Our plants and suppliers say they have capacity, our plan assumes they do, and then they don’t deliver. How do we validate capacity before we plan against it?

Why it happens: Capacity is stated as a theoretical rate (units per shift at nameplate) and the plan treats it as available. Actual capacity is lower by changeovers, maintenance, absenteeism, quality loss, and the fact that the same line is committed to another product family. A supplier states capacity in the same optimistic terms, because saying yes wins the business. The plan is feasible on the stated number and infeasible on the real one, and the shortfall appears as missed schedules that everyone blames on execution.

Do this first: For each constrained resource, compare demonstrated capacity (what it actually produced over the last quarter, by week) to the number in the planning system. Use demonstrated capacity in the plan. For suppliers, ask for demonstrated rate on your parts and the share of their line you are counting on, and validate with a visit on the ones that matter. Then track capacity attainment (actual versus planned) as a standing metric so the number in the plan drifts toward reality instead of away from it.

What good looks like: The plan is built on demonstrated, not theoretical, capacity, supplier capacity is validated on critical parts, and capacity attainment is reviewed monthly with owners for the gaps.

PlanningPlanning Systemscascading outages

Q: A small number of parts cause most of our shortages and expedites, and they get the same planning attention as everything else. How should we manage them differently?

Why it happens: Planning systems treat every item the same way, and planners spread their attention across thousands of items. But shortages concentrate: a few dozen parts with long lead times, single sources, or high demand variability drive most of the line-down events and most of the premium freight. Without a separate governance for those parts, problems are found at the same speed as problems on the tail, which is too slow.

Do this first: Identify the constraint parts with data: highest expedite spend, highest shortage frequency, longest lead time, single source, or all four. It is usually under a hundred items. Put them on a weekly review with Planning, Procurement, and the plant or DC, looking at projected inventory against demand for the next eight weeks, open PO status, and supplier confirmations. Give each part an owner. Set the rule that any projected shortage on the list triggers action that week, and track how many shortages were prevented versus how many were expedited.

What good looks like: Constraint parts are reviewed weekly with projected coverage, shortages are seen three or more weeks out rather than three days out, and expedite spend on the list falls quarter over quarter.

Grocery and short-shelf-life: Constraint governance here is as much about date codes and rotation as about quantity. Add remaining shelf life to the weekly view.

PlanningPlanning SystemsGrocery & Foodexpediting

Q: We expedite something every day. It works, so nobody questions it, but premium freight is eating our margin. How do we get out of the expedite cycle?

Why it happens: Expediting succeeds locally and fails globally. Each expedite saves a customer order or a production run, so it is rewarded. But every expedite pulls capacity, labor, and attention from the plan, which creates the next shortage, which needs the next expedite. Meanwhile the root causes (bad lead times, unvalidated capacity, unreliable suppliers, forecast bias) are never fixed because the fire is always out by the time anyone could look at it. The organization gets very good at expediting and never gets good at planning.

Do this first: Log every expedite for 30 days with its cost and its cause: late supplier, plan error, demand spike, data error, capacity miss. The distribution will surprise you; it is rarely demand. Then attack the top cause with the relevant note in this collection (lead times, capacity, constraint governance, supplier recovery). In parallel, put a budget and an approval on expedites so the cost is visible and the decision is deliberate. Report expedite spend against cause weekly until it falls.

What good looks like: Expedite spend is a small, budgeted line with known causes, the plan is trusted enough that people wait for it, and the planners who used to be heroes are now the people who prevent the fire.

PlanningStabilize & RecoveryPlanning Systemsexpediting

Q: Our safety stock is based on demand variability only. Supplier lead time varies by weeks and we get caught every time. How do we plan for it?

Why it happens: Most planning setups model demand uncertainty and treat supply as deterministic: the lead time is a number, and the order will arrive on that date. Real lead times have a distribution, and for many suppliers the spread is wider than the demand spread. When a shipment is two weeks late, the demand-based safety stock covers two weeks of demand variability, not two weeks of demand. The stockout follows, and the planner responds by inflating the lead time parameter, which adds inventory everywhere without protecting anything specifically.

Do this first: Measure lead time variability by supplier and item from receipt data (see the note on lead times). Then, where the planning system supports it, include supply variability in the safety stock formula; where it doesn’t, segment items by supply reliability and set differentiated safety stock rules by segment. Separately, work on reducing the variability: confirmed ship dates, ASN visibility, and supplier recovery discipline reduce the spread, which reduces the buffer needed. Review the top 50 items by lead time spread with Procurement quarterly.

What good looks like: Safety stock reflects both demand and supply variability, the suppliers with the widest spread are on an improvement plan, and inventory is positioned by measured risk rather than by a uniform rule.

Automotive: On sequenced or line-critical parts, lead time variability is a line-down risk, not an inventory risk. Combine the modeling with constraint-part governance and a recovery plan.

PlanningPlanning SystemsAutomotiveCPGexcess & stockouts

Q: Planners override the system’s recommendations most of the time. Are they fixing the system or hiding what’s wrong with it?

Why it happens: Usually both, and the system can’t tell the difference. An experienced planner knows the supplier is late, the forecast is wrong on promotions, the MOQ is negotiable, so they override. Each override is correct and each one prevents the underlying parameter or data problem from ever being fixed, because the outcome looked fine. When the planner leaves, the overrides leave with them, and the system’s recommendations are exposed as being as wrong as they always were.

Do this first: Require a reason code on planning overrides and report override rate by planner, by item, and by reason. Review the top reasons monthly with the planners. Every reason that repeats is a system fix: a lead time to update, a forecast model to adjust, a parameter to change, a supplier issue to escalate. Make those fixes, then watch the override rate on that reason fall. Keep planner judgment where it is genuinely judgment, and get it out of the places where it is compensation for a known defect.

What good looks like: Override rate is under 20% and falling, every override has a reason, and the monthly review is a list of system improvements rather than a list of exceptions.

Grocery: High-frequency replenishment and promotions make override rates naturally higher. The target isn’t zero; it’s that the reasons are known and the repeatable ones are encoded.

PlanningPlanning SystemsGrocery & Foodoverride behavior

Q: We plan a production or fulfillment sequence and the floor runs a different one. We don’t know how often or what it costs. Should we measure it?

Why it happens: Sequence adherence is invisible in most systems because the schedule is a plan and the execution is a set of transactions, and nobody joins them. The floor resequences for good reasons (a material shortage, a changeover saving, a hot order) and bad ones (habit, convenience). Each deviation ripples: the next job’s material isn’t staged, the downstream cell waits, the promised date slips. Without a measure, the plan is treated as advisory and the schedule’s credibility erodes.

Do this first: Compare planned sequence to actual sequence for a week, by line or by cell, and count the deviations with a reason. Then classify: deviations caused by the plan being wrong (material wasn’t there, capacity was overstated) and deviations caused by execution choosing differently. Fix the plan problems in the planning system; address the execution problems with the supervisors and a simple rule about who can resequence and how it is recorded. Publish sequence adherence weekly alongside schedule attainment.

What good looks like: Sequence adherence is above 90% on planned lines, every deviation has a reason, and the planning team learns from the deviations to make the next schedule more feasible.

Medical and automotive: In regulated manufacturing and sequenced supply, an unrecorded resequence is a traceability and compliance issue as well as an efficiency one. Make the resequence transaction mandatory.

PlanningPlanning SystemsMedical & RegulatedAutomotiveKPI decline

Q: Defects jump every time a supplier changes something (a material, a process, a site) and we find out after the product is in our building. How do we get ahead of supplier changes?

Why it happens: Suppliers change things without telling you because nobody told them they had to, or because the requirement is buried in a quality agreement they signed years ago. A new resin, a new sub-tier, a new plant, a new inspection sampling plan: each is invisible until your incoming inspection or your customer catches the difference. By then the product is spread across inventory and orders, and the containment costs more than the change would have.

Do this first: Put a change notification requirement in the quality agreement with a definition of what counts as a change and a lead time (30 to 90 days is typical), and enforce it through the scorecard. On your side, build a supplier change register: what changed, when, who approved it, what verification was done. Tie incoming inspection to the register so a changed part gets a tighter sampling plan for its first lots. When a defect spike happens, check the register first; if there is no recorded change, the conversation with the supplier starts there.

What good looks like: Supplier changes are notified before they ship, verified on receipt with a defined plan, and traceable in a register, and defect spikes are explained rather than investigated from scratch.

Food and medical: Change control here is a regulatory requirement, not a best practice. Make sure the register can produce evidence for an audit or a recall investigation.

Implement & Go-LiveStabilize & RecoveryQuality SystemsGrocery & FoodMedical & Regulatedsupplier noncompliance

Q: Our corrective and preventive action backlog keeps growing. Items are overdue, auditors are asking questions, and the team is demoralized. How do we get it under control?

Why it happens: CAPAs are opened faster than they are closed because opening one is easy and closing one requires root cause, an action, verification of effectiveness, and a sign-off. When every nonconformance becomes a CAPA regardless of severity, the queue fills with items that never needed the full process, and the significant ones get the same attention as the trivial ones. Owners are assigned but not resourced, due dates are set by policy rather than by the work, and the metric everyone watches is the backlog count, which drives closure on paper rather than in practice.

Do this first: Triage the backlog by risk: which open items relate to a safety, regulatory, or customer-impact issue. Work those first with real resources. For the rest, apply a risk-based rule for what needs a full CAPA versus a simpler correction, and close or downgrade the items that don’t meet the bar, with documented rationale. Then change the intake: a nonconformance is escalated to a CAPA by a defined risk criterion, not by default. Track age and effectiveness verification, not just count, and review the top ten open items weekly with the people who own the resources.

What good looks like: The CAPA backlog is short and risk-ranked, new CAPAs are opened by criterion, closure includes verified effectiveness, and the auditor’s question is answered by the register rather than by explanation.

Aerospace and medical: AS9100 and ISO 13485 auditors care about the effectiveness check more than the closure date. Design the process around that and the backlog stops being the problem.

Stabilize & RecoveryImplement & Go-LiveQuality SystemsAerospaceMedical & RegulatedKPI decline

Q: Inspection records, certs, and test results are captured differently by shift and by site. When a customer or auditor asks for evidence, it takes days to assemble. How do we standardize?

Why it happens: Evidence capture was designed around forms, and forms get filled in differently by different people. One shift scans the cert, another files the paper, a third writes the lot number on the traveler. The quality system has fields that aren’t mandatory, so they are skipped under pressure. When the request comes (a customer complaint, an audit, a recall inquiry), the evidence exists somewhere but not in one place, and the assembly is a manual search across systems, cabinets, and inboxes.

Do this first: Define the evidence set for each product family: what records must exist, in what form, linked to what identifier (lot, serial, work order, receipt). Make the mandatory fields mandatory in the system, and block the next step until they are complete. Give the floor the tools to capture at the point of work (scan, photo, device upload) so the record is created once. Then test the retrieval: pick a lot at random and time how long it takes to assemble the complete evidence. Repeat monthly until it is minutes.

What good looks like: Every unit of product has a complete, linked evidence set, capture is enforced by the system rather than by inspection, and any lot’s history can be assembled in minutes by anyone in Quality.

Stabilize & RecoveryImplement & Go-LiveQuality SystemsAerospaceMedical & Regulatedtraceability gaps

Q: Our process controls hold when volume is normal and collapse at peak or during a shortage. Inspection is skipped, holds are released early, and defects escape. How do we make controls that survive pressure?

Why it happens: A control that depends on a person choosing to follow it will be skipped when the person has a stronger incentive not to. Under pressure, the incentive is throughput. The inspection that takes ten minutes is sampled instead of performed, the hold that blocks a shipment is released by a supervisor who needs the truck to leave, the deviation is approved after the fact. Each choice is locally rational and each one erodes the control until it exists only on paper.

Do this first: Classify your controls into system-enforced (the WMS won’t ship a lot on hold, the MES won’t advance a job without the inspection result) and person-enforced (someone checks, someone signs). Convert the highest-risk person-enforced controls into system-enforced ones; that removes the choice. For the ones that must stay person-enforced, measure adherence (skip rate, early release rate, retrospective approvals) and review it weekly, so pressure shows up as a number before it shows up as an escape. Give the people under pressure a legitimate fast path (expedited inspection, a defined release authority) so the shortcut isn’t the only option.

What good looks like: The critical controls are enforced by the system and can’t be skipped, adherence to the remaining controls is measured, and peak season is planned with the inspection capacity it needs rather than with the hope that people will cope.

Implement & Go-LiveStabilize & RecoveryQuality SystemsMedical & RegulatedAerospaceIndustrialexception overload

Q: Every time something goes wrong on the floor, the fix leaves a gap in the quality record, and the auditor finds it months later. How do we make exception handling audit-safe?

Why it happens: The standard process is documented and controlled; the exception process is improvised. A mislabeled lot is relabeled by hand, a rejected receipt is accepted with a phone-call approval, a rework is done without a rework order, an adjustment is made in the WMS without a deviation record. The product is fine, the operation moves on, and the quality record has a hole where the exception happened. The auditor doesn’t care that the product was fine; they care that you can’t prove it.

Do this first: List the exception types that happen on the floor (there are usually ten or so) and give each one a documented path that produces a record: a deviation, a rework order, a concession, a controlled relabel. Make the record part of doing the exception, not a form filled in afterward: the system creates the deviation when the operator selects the exception reason. Then audit yourself: pull last month’s WMS adjustments and inventory moves and check that each one with a quality implication has a matching record. The ones that don’t are your training and design list.

What good looks like: Every exception produces a quality record as a by-product of handling it, internal audits find no undocumented exceptions, and the floor sees the exception path as the fast path because it is designed to be.

Stabilize & RecoveryQuality SystemsMedical & RegulatedAerospacetraceability gaps

Q: Our conveyor and sortation system jams at the same points every day and the integrator says it’s the product. Is it, and what do we do?

Why it happens: Sometimes it is the product (a carton outside spec, a bad label placement), but recurring jams at the same point are usually a logic or configuration issue: a merge that releases too aggressively, a scanner tolerance set for a mix that has since changed, a divert timing that was tuned for a case size you no longer run, or accumulation zones that are too short for the current wave profile. The integrator tuned the system for the design case. The operation is running a different case.

Do this first: Log jams by location, time, and product characteristic for a week. If jams cluster at a location, walk it with the controls engineer while it is running and watch what the logic does under load. Adjust one parameter at a time (merge release, gap, scanner tolerance, divert timing) and measure. If jams cluster by product, tighten the carton and label specification at the source and enforce it at induction. In parallel, make sure the WCS or WES is telling you the jam rate and the recovery time as a metric, because the goal is throughput, not zero jams.

What good looks like: Jam rate per thousand cartons is tracked and trending down, recovery time is under a defined target because the recovery procedure is standard work, and the controls parameters are reviewed whenever the product mix or wave profile changes.

Stabilize & RecoveryImplement & Go-LiveAutomationGrocery & FoodCPGcascading outages

Q: The automated system handles the normal flow beautifully and then dumps every exception on a person with a clipboard. The exception lanes are the bottleneck. How do we automate the exceptions too?

Why it happens: Automation projects are scoped on the main flow and the exceptions are left to “manual handling” because they seemed rare. In production they are 5–15% of volume: no-reads, weight mismatches, overflows, unexpected items, missing labels. Each one exits the system to a station where a person figures out what happened, corrects it, and reintroduces it, with no system support and no measurement. The station backs up, product ages, and the automation’s throughput is capped by a person with a clipboard.

Do this first: Measure exception volume by type at each exit point. Then, for the top three types, design a system-supported resolution: a screen that shows the operator what the system knows, a one-scan re-induction, an automatic re-label, a rule that resolves the common case without a person. Where an exception can be prevented at the source (label quality, weight tolerance, carton spec), fix the source. Staff the exception station from the measured volume, and give it a queue with age visibility so it is managed like any other work.

What good looks like: Exception rate is under 3% and each type has a defined resolution path, the exception station clears continuously rather than at end of shift, and the automation team reviews the top exception type weekly to design it out.

Implement & Go-LiveStabilize & RecoveryAutomationGrocery & Foodexception overload

Q: Our new automated system is rated for far more than we get out of it, and the constraint is the people feeding it. How do we get induction to keep up?

Why it happens: The system was sized on its downstream rate and induction was assumed. In practice, induction is where every upstream problem arrives: product not staged, cartons not labeled, totes not available, a screen that needs four scans, an operator who has to walk to get the next pallet. The automation waits on the human, and the human is waiting on the process. The rated throughput is real; the feed is what was never designed.

Do this first: Time the induction cycle by element for a shift: walk, retrieve, scan, place, confirm. The elements that aren’t the value-adding “place” are your targets. Stage product to the induction point ahead of need, reduce the scan sequence to the minimum the system needs, and give induction its own replenishment labor so operators never leave the station. Then look at the upstream release logic: is work arriving at induction in a sequence that requires the operator to sort or search? Fix the release, not the operator. Measure induction rate per station per hour and compare it to the downstream rate; the gap is your improvement target.

What good looks like: Induction rate is within 10% of the system’s downstream rate, operators stay at the station, and the induction area is fed by a scheduled flow rather than by whoever is free.

Industrial and automotive: Heavy, irregular, or kitted product needs induction fixtures and handling aids designed with the operator. The bottleneck here is often ergonomics, not process.

Implement & Go-LiveStabilize & RecoveryAutomationIndustrialAutomotivehandoff failures

Q: Our goods-to-person or pick-to-light system starves because replenishment into it can’t keep pace with picking out of it. How do we balance the flow?

Why it happens: Automated picking is only as fast as the slowest input, and replenishment into the automation is usually a manual process planned around yesterday’s picking. The system runs down its inventory of a fast mover mid-wave, requests replenishment, and waits. Underneath: replenishment is triggered reactively rather than forecast from the wave, the decant or induction into the automation is under-staffed relative to pick rate, and the item mix in the automation was set at design and never re-slotted for velocity.

Do this first: Forecast replenishment from the released waves rather than from the system’s minimums: when the wave is built, calculate demand by item against inventory in the automation and release replenishment tasks ahead of the picks. Staff the decant or induction stations from that forecast, by hour. Then review which items are in the automation at all: fast movers that exhaust every wave may need more locations or a case-pick path outside the system, and slow movers may be occupying space that fast movers need. Track starvation events (system waiting on replenishment) as the key metric.

What good looks like: Starvation events are rare and measured, replenishment runs ahead of picking, and the item mix and location count in the automation are reviewed monthly against pick velocity.

Implement & Go-LiveStabilize & RecoveryAutomationshort picks

Q: We installed an AS/RS to raise throughput and the numbers are about where they were before. Where is the capacity going?

Why it happens: An AS/RS delivers its rated throughput only when the storage strategy, the task mix, and the interfaces are tuned to the actual operation. Common losses: product slotted by size rather than velocity, so the cranes travel to the far end for every fast mover; retrieval and storage tasks not interleaved, so cranes run half their cycles empty; the WMS releasing tasks in an order that creates contention on one aisle while the others idle; and the input and output conveyors buffering too little, so the cranes wait on the conveyor and the conveyor waits on the cranes. Each loss is 10–15%; together they eat the improvement.

Do this first: Get the crane utilization and cycle data from the WCS: task counts by type, travel distance by task, dual-cycle rate, wait time on input and output. That tells you where the capacity is going. Re-slot by velocity within the aisles, turn on task interleaving, balance the workload across aisles in the WMS release logic, and size the buffers to the crane’s rate. Change one thing at a time and measure throughput over a full shift. If the integrator tuned the system on the design data, ask them to re-tune on six months of actual.

What good looks like: Dual-cycle rate is above 70%, aisle utilization is balanced, the cranes rarely wait on conveyors, and throughput is within 15% of the rated number during peak hours.

Grocery and automotive: Temperature zones and sequenced outbound add constraints to slotting and release. Tune inside them; don’t assume the system will find the optimum on its own.

Stabilize & RecoveryAutomationGrocery & FoodAutomotiveKPI decline

Q: Since we automated, small data and labeling errors that used to be caught by people now cause large problems. Is automation making us worse?

Why it happens: Automation removes the person who used to quietly fix things. A picker who saw a wrong label would correct it; a sorter that sees a wrong label sends the carton to the wrong destination at 6,000 cartons an hour. A receiver who noticed a pack quantity mismatch would count it; an automated decant that trusts the ASN loads the wrong quantity into a tote and the error propagates to every order picked from it. Automation didn’t create the defects. It made their cost visible by executing them faithfully.

Do this first: Trace the last twenty automation incidents to their origin. Most will start upstream: a supplier label, an item master dimension, a pack quantity, an order that was keyed wrong. That list is your validation backlog. Put hard checks at the boundary before product or data enters the automation: label verification at induction, dimension and weight validation against the item master, quantity confirmation on decant, order validation before release. Reject at the boundary and route to a resolution station, rather than letting the defect in. Then fix the top upstream sources with their owners (supplier, master data, order entry), because the validation is a containment, not a cure.

What good looks like: Defects are caught at the automation boundary rather than inside it, upstream data quality is measured by the automation’s rejection rate, and the people who used to fix things quietly are now fixing the sources.

Medical and industrial: A defect that propagates through automation in a regulated flow can require a lot-level recall of everything the automation touched. Validation at the boundary is a compliance control here, not an efficiency one.

Stabilize & RecoveryImplement & Go-LiveAutomationMedical & RegulatedIndustrialdata trust loss

No notes match that combination

Try fewer filters or a broader search term. If the problem you are dealing with isn’t here yet, tell us about it and we will write the note.

Describe your situation

Recognize your operation in one of these?

Send us the facts (the exception report, the volumes, the cutoffs you are missing) and we will tell you where the constraint is. No deck, no sales cycle, just a straight read on what to fix first.

When to Call Us