| .. | ||
| best.pt | ||
| README.md | ||
| shelf-products-detection-train-yolov8n.ipynb | ||
Phase 1 — Shelf Product Detection with YOLOv8n
Pipeline Documentation · SKU110K Dataset
Overview
This document covers Phase 1 of a two-phase retail shelf analysis pipeline. The goal of Phase 1 is to detect all products visible on retail shelves using a fine-tuned YOLOv8n object detection model trained on the SKU110K dataset. The output (bounding boxes around each product) feeds directly into Phase 2, which classifies each detected product.
Shelf Image → [Phase 1: YOLOv8n Detection] → Bounding Boxes → [Phase 2: Classifier] → Product Labels
Environment
| Item | Details |
|---|---|
| Platform | Kaggle Notebooks (GPU) |
| GPU | Required (auto-selected with batch=-1) |
| Framework | Ultralytics YOLOv8 |
| Language | Python 3 |
Prerequisites
1. Install Dependencies
pip install -q ultralytics
This installs the Ultralytics library, which includes YOLOv8 and all required dependencies (PyTorch, OpenCV, etc.).
Dataset — SKU110K
What It Is
SKU110K is the primary dataset used for training. It contains:
- 11,762 shelf images capturing dense retail environments
- ~1.2 million annotated products (bounding boxes)
- Designed for dense object detection in challenging, cluttered shelf scenarios
| Resource | Link |
|---|---|
| GitHub (paper + info) | https://github.com/eg4000/SKU110K_CVPR19 |
| Kaggle (download) | https://www.kaggle.com/datasets/thedatasith/sku110k-annotations |
Alternative Datasets (for future reference)
If you need to augment or swap the dataset, consider:
- Grocery Store Dataset (Grozi-120) — 120 grocery product classes
- WebMarket — web-scraped retail product images
- RPC (Retail Product Checkout) — checkout-counter product recognition
Step-by-Step Reproduction Guide
Step 1 — Locate the Dataset Folder
After adding the SKU110K dataset to your Kaggle notebook, run the following to find and verify the dataset structure:
import os
BASE_PATH = "/kaggle/input/datasets/thedatasith/sku110k-annotations"
DATASET_FOLDER = None
# Find SKU110K_fixed folder dynamically
for item in os.listdir(BASE_PATH):
if "SKU110K" in item:
DATASET_FOLDER = os.path.join(BASE_PATH, item)
break
print("📁 Dataset folder:", DATASET_FOLDER)
# Print directory tree (2 levels deep)
for root, dirs, files in os.walk(DATASET_FOLDER):
level = root.replace(DATASET_FOLDER, '').count(os.sep)
indent = ' ' * 2 * level
print(f"{indent}📁 {os.path.basename(root)}/")
for f in files[:5]:
print(f"{indent} 📄 {f}")
if level >= 2:
break
Why: The folder name inside the dataset archive contains SKU110K but may have a suffix (e.g., SKU110K_fixed), so we detect it dynamically rather than hardcoding it.
Step 2 — Inspect the YAML Configuration
The dataset comes with a data_kaggle.yaml file that defines paths and class names. Read it to verify its contents before making any changes:
import yaml
YAML_PATH = BASE_PATH + "/data_kaggle.yaml"
with open(YAML_PATH, "r") as f:
data = yaml.safe_load(f)
print("Classes:", data.get("names"))
print("Number of classes:", len(data.get("names", [])))
print("Train path:", data.get("train"))
print("Val path:", data.get("val"))
Step 3 — Fix the YAML Paths
The paths in the original YAML file point to locations that don't match the Kaggle input filesystem. You must override them with the correct absolute paths before training:
import yaml
YAML_PATH = BASE_PATH + "/data_kaggle.yaml"
with open(YAML_PATH, "r") as f:
data = yaml.safe_load(f)
# Override with correct Kaggle paths
data["train"] = DATASET_FOLDER + "/images/train"
data["val"] = DATASET_FOLDER + "/images/val"
# Save the corrected YAML to the working directory
FIXED_YAML_PATH = "/kaggle/working/fixed_data.yaml"
with open(FIXED_YAML_PATH, "w") as f:
yaml.dump(data, f)
print("✅ Fixed YAML saved at:", FIXED_YAML_PATH)
print(data)
Important: Always write the fixed YAML to
/kaggle/working/, not back to the input directory (which is read-only on Kaggle).
Step 4 — Train YOLOv8n
Verify GPU availability and launch training:
import torch
import numpy
from ultralytics import YOLO
# Verify GPU
print("GPU:", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "None")
print("CUDA Capability:", torch.cuda.get_device_capability(0))
# Load pretrained YOLOv8 nano model
model = YOLO('yolov8n.pt')
# Train
model.train(
device='0', # GPU index
data=FIXED_YAML_PATH, # Path to the fixed dataset config
epochs=30, # Total training epochs
imgsz=800, # Input image size (px)
batch=-1, # Auto-select largest batch that fits in GPU memory
workers=4, # Dataloader worker threads
optimizer="AdamW", # Optimizer
lr0=0.002, # Initial learning rate
patience=10, # Early stopping: stop if no improvement for 10 epochs
save_period=5, # Save a checkpoint every 5 epochs
project="kaggle/working/runs", # Output directory
name="exp", # Experiment name (auto-increments: exp, exp2, ...)
save=True # Save final weights
)
Training Hyperparameters Explained
| Parameter | Value | Rationale |
|---|---|---|
imgsz |
800 | Higher resolution improves detection of small/dense products |
batch |
-1 | Auto-batch maximizes GPU utilization without OOM errors |
optimizer |
AdamW | Better generalization than SGD for this type of task |
lr0 |
0.002 | Starting learning rate |
patience |
10 | Stops training early if validation loss plateaus |
save_period |
5 | Checkpoint every 5 epochs for resumability |
epochs |
30 | Total epochs (may stop earlier due to patience) |
Step 5 — Download Trained Weights
After training completes, package the experiment output folder as a ZIP file and create a download link:
import shutil
from IPython.display import FileLink
# Update this path to match your actual experiment folder name
folder_path = '/kaggle/working/runs/detect/kaggle/working/runs/exp-6'
zip_path = '/kaggle/working/exp-6.zip'
shutil.make_archive('/kaggle/working/exp-6', 'zip', folder_path)
print("Weights saved to:", zip_path)
FileLink(zip_path)
Note: The experiment folder name auto-increments (e.g.,
exp,exp-2,exp-6). Check/kaggle/working/runs/detect/to find your folder name before running this cell.
The ZIP will contain:
weights/best.pt— best checkpoint (highest validation mAP)weights/last.pt— final checkpoint (used for resuming)- Training curves, confusion matrix, and metrics logs
Step 6 — Resume an Interrupted Training Run
If training is interrupted (e.g., Kaggle session timeout), you can resume from the last saved checkpoint. Upload last.pt as a Kaggle dataset input first, then:
from ultralytics import YOLO
model = YOLO("/kaggle/input/weights/last.pt")
model.train(resume=True)
resume=Trueautomatically picks up the training configuration, optimizer state, and epoch counter from the checkpoint.
Output Files
After a successful training run, the experiment folder contains:
exp/
├── weights/
│ ├── best.pt ← Use this for inference and Phase 2 input
│ └── last.pt ← Use this to resume training
├── results.csv ← Per-epoch metrics (loss, mAP, precision, recall)
├── results.png ← Training curves plot
├── confusion_matrix.png
├── val_batch*.jpg ← Validation batch visualizations
└── args.yaml ← Full training config (for reproducibility)
Using the Trained Model for Inference (Phase 2 Input)
After training, use best.pt to detect products in new shelf images:
from ultralytics import YOLO
model = YOLO("best.pt")
results = model.predict(source="shelf_image.jpg", conf=0.25, iou=0.45)
for result in results:
boxes = result.boxes.xyxy # Bounding boxes [x1, y1, x2, y2]
confs = result.boxes.conf # Confidence scores
# Pass boxes to Phase 2 classifier
Troubleshooting
| Problem | Likely Cause | Fix |
|---|---|---|
DATASET_FOLDER is None |
Folder name doesn't contain "SKU110K" | Print os.listdir(BASE_PATH) and update the if condition |
| YAML path errors during training | Paths still point to original locations | Re-run Step 3 and confirm FIXED_YAML_PATH exists |
| CUDA out of memory | Batch size too large | Set batch to a fixed value (e.g., 8 or 16) instead of -1 |
| Training not resuming | Wrong path to last.pt |
Upload last.pt as a Kaggle dataset, verify the input path |
| Experiment folder not found | Auto-increment changed the name | List /kaggle/working/runs/detect/ to find the correct name |
Pipeline Context
┌─────────────────────────────────────────────────────────────────────┐
│ Full Pipeline │
│ │
│ Input Image │
│ │ │
│ ▼ │
│ ┌─────────────────────────────────┐ │
│ │ PHASE 1 — Detection │ ← This notebook │
│ │ YOLOv8n trained on SKU110K │ │
│ │ Output: bounding boxes │ │
│ └─────────────────────────────────┘ │
│ │ │
│ │ Cropped product regions │
│ ▼ │
│ ┌─────────────────────────────────┐ │
│ │ PHASE 2 — Classification │ ← Next phase │
│ │ Classify each detected product │ │
│ │ Output: product labels │ │
│ └─────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
Phase 1 produces bounding boxes for every product on the shelf. Each cropped region is then passed to the Phase 2 classifier for product identification.
Documentation generated from shelf-products-detection-train-yolov8n.ipynb