Give bounded context.
Provide the mission objective, approved evidence, constraints, and required output rather than asking for a generic solution.
Build a maritime delay decision system that reconciles unstable operational evidence and exposes financial risk before the decision window closes.
The lab permits agent-assisted investigation and implementation. Strong work records material agent recommendations, checks them against authorized evidence, and documents what the learner accepted, changed, rejected, or left unresolved.
Provide the mission objective, approved evidence, constraints, and required output rather than asking for a generic solution.
Require multiple hypotheses, failure modes, and disconfirming evidence before choosing an intervention.
Check source IDs, calculations, code behavior, policy constraints, and unsupported causal language.
Record why the final decision follows from the evidence and what would cause it to change.
The Core Labs isolate individual failure modes. This mission forces several of them to coexist under one customer timeline, one evidence room, and one final deployment recommendation.
Should the shipment-risk workflow move from pilot to broader operational use, and under what data and authority conditions?
Shipment, vessel, port, booking, and cost sources disagree on identity, timing, and exposure while operators still need a decision before the data is perfect.
Scale, narrow, revise, continue the pilot, or stop — with an explicit identity policy, failure behavior, adoption evidence, and value case.
The simulated freight operator coordinates maritime shipments for industrial customers. A dispatcher may need 42 minutes to determine whether a vessel delay threatens a contractual delivery window and creates financial exposure.
The company has already assembled a dashboard prototype. It displays vessel positions and port events, but users continue to investigate in spreadsheets, email, and carrier portals because the system cannot explain disagreements between sources or show which data is stale.
Your mission: design the minimum deployable workflow that reduces median investigation time below ten minutes while preserving uncertainty and human authority.
The design must resolve real differences in incentives rather than averaging them into a generic requirement list.
Wants one risk queue and a measurable reduction in investigation time.
Knows the exceptions and fears the system will hide uncertainty.
Wants a canonical model, stable contracts, and no direct load on production systems.
Requires explainable exposure calculations before discussing avoided cost.
The guided workspace now includes downloadable interviews, operational documents, dirty CSV/JSON evidence, templates, a notebook, progressive hints, and a staged complication.
Dispatchers describe five different definitions of “delay” depending on the downstream decision.
Near-real-time positions with occasional gaps, duplicate pings, and changing vessel names.
Estimated and actual arrival/departure events using a different port code system.
Customer bookings and contractual windows, refreshed twice daily from a spreadsheet export.
Demurrage and operational exposure logic maintained by one analyst with undocumented overrides.
Messages showing how teams currently resolve disagreements between feeds and customer commitments.
A partial mapping among IMO numbers, internal vessel IDs, names, and carrier references.
Customer data cannot appear in the first pilot; production access requires separate review.
Progress is stored locally in your browser. The sequence intentionally moves from decision and evidence to architecture, release, adoption, and value.
The mission is divided into reviewable checkpoints so the final solution cannot be reverse-engineered from the complication.
Submit the decision statement, stakeholder map, current workflow, baseline, constraints, and unresolved questions.
Gate: problem framingSubmit the minimum deployable workflow, canonical model, source contracts, architecture, and risk register.
Gate: technical judgmentSubmit the working slice, failure behavior, rollout plan, training, rollback, and measurement instrumentation.
Gate: operational credibilityRevise after the complication and pilot evidence, then recommend scale, revision, continued pilot, or stop.
Gate: value judgmentThe exact complication is intentionally withheld here. In the guided workspace, it unlocks only after you complete the working-slice stage and preserve your original assumptions.
After revising the design, learners receive a complete pilot package covering usage, investigation time, overrides, false urgency, data quality, operator feedback, and executive pressure to scale.
Inspect freshness, identity confidence, exposure completeness, failure modes, and the difference between median and tail behavior.
Determine whether usage reflects genuine workflow adoption or continued dependence on spreadsheets and manual verification.
Evaluate investigation speed, override reasons, exception handling, and whether the system changes the intended decision.
Recommend scale, revision, narrowing, continuation, or stop without confusing executive enthusiasm with proven value.
The strongest submissions are not the largest systems. They are coherent, explicit about uncertainty, and tied to the operational decision.
| Dimension | What strong work demonstrates | Weight |
|---|---|---|
| Problem frame | Decision statement, workflow map, stakeholders, baseline, assumptions, and scope. | 15% |
| Data and identity design | Canonical model, source profile, matching rules, confidence states, and contracts. | 20% |
| Deployable workflow | Architecture or implementation covering interface, services, access, and failures. | 20% |
| Pilot and adoption | Users, training, support, overrides, escalation, rollout, and rollback. | 15% |
| Measurement | Technical, adoption, operational, and financial measures tied to the baseline. | 15% |
| Executive recommendation | Scale, revise, or stop—with evidence, uncertainty, risks, and next actions. | 15% |