GitHub - ArgoHA/D-FINE-seg: D-FINE-seg Object Detection and Segmentation Framework (Train, Export, Inference)
Pangram verdict · v3.3
We believe that this document is a mix of AI-generated, and human-written content
AI likelihood · overall
MixedArticle text · 1,197 words · 5 segments analyzed
Real-Time Object Detection, Instance and Semantic Segmentation Quick Start • Usage • Export • Inference • Benchmarks • Video Tutorial • Colab
D-FINE-seg is a framework for real-time object detection, instance segmentation, and semantic segmentation - one codebase, one config flag (task: detect | segment | sem_seg), five model sizes (N -> X).
End-to-end workflow - dataset prep -> training (DDP, EMA, AMP, mosaic) -> export (ONNX, TensorRT, OpenVINO, CoreML, LiteRT) -> benchmarked multi-backend inference Accuracy - on Cityscapes, beats YOLO26 and RF-DETR on detection & instance-seg F1 and leads mIoU on semantic segmentation, at real-time latency with 2-3x fewer params; also higher F1 than YOLO26 on TACO and VisDrone (TensorRT FP16, end-to-end protocol) Paper - D-FINE-seg: Object Detection and Instance Segmentation Framework with Multi-Backend Deployment Not a fork: the detection core follows the D-FINE paper; segmentation heads, training, export and inference are implemented from scratch.
One frame, three tasks, one config flag:
Full tables below
Highlights
Instance segmentation head (task: segment) - lightweight mask head on top of D-FINE's HybridEncoder PAN outputs: stride 8/16/32 features fused to 1/4 resolution, then a dot-product between per-query mask embeddings (3-layer MLP) and the shared mask features yields per-instance masks
Semantic segmentation head (task: sem_seg) - reuses the pretrained instance-seg mask fuser on full-frame features, followed by a small conv neck and 1x1 classifier: no queries, no NMS Mask-aware training - box-cropped BCE + Dice mask losses (instance seg) and CE + multi-class soft Dice with ignore_index (semantic seg), mask supervision inside contrastive denoising, and Dice + sigmoid-focal mask costs in the Hungarian matcher - all train-time only, zero inference cost COCO-pretrained weights for detection and instance segmentation, auto-downloaded on first use - fine-tuning starts from a trained mask decoder, not from scratch Multi-channel inputs - train on RGB + thermal / depth / NIR stacks (4-channel .npy), not just RGB Modern training stack - Muon optimizer, DDP, EMA, mosaic + affine augs, OneCycle, early stopping, WandB Beyond the model - ByteTrack tracking, SAM3 auto-labeling, Gradio demo, INT8 quantization (OpenVINO / CoreML / LiteRT)
Quick Start Installation git clone https://github.com/ArgoHA/D-FINE-seg.git cd D-FINE-seg uv sync This creates a .venv/ with all dependencies pinned by uv.lock. Activate it with source .venv/bin/activate, or run anything via uv run ... (the Makefile already does this). Pretrained weights are auto-downloaded from Hugging Face into pretrained/ on first use, so no manual setup is needed. To download manually instead, grab dfine_<size>_<dataset>.pt (size ∈ {n, s, m, l, x}, dataset ∈ {coco, obj2coco}) and place it in pretrained/. Segmentation weights are also available in the Hugging Face model card. Prepare Your Data Two annotation formats are supported: YOLO (default) and COCO JSON. Semantic segmentation uses PNG masks instead (see below). YOLO format (default) data/dataset/ ├── images/ # all images: .jpg, .png, etc. (
.npy for multi-channel - see below) └── labels/ # all labels: one .txt per image (same filename stem) Detection labels: class_id xc yc w h (normalized) Segmentation labels: class_id x1 y1 x2 y2 ... xN yN (normalized polygon coordinates) Input types & channel order: 3-channel .jpg/.png (BGR, read via cv2.imread), 3-channel .npy (RGB, read via np.load), or 4-channel .npy (RGB+extras, e.g. RGB+thermal). Semantic segmentation masks (task: sem_seg) data/dataset/ ├── images/ # same as YOLO layout └── labels/ # one single-channel uint8 .png per image (same stem), pixel value = class id Every pixel gets a class from label_to_name (background included). Pixels with value train.sem_seg.ignore_index (default 255) are excluded from loss and metrics and during inference 255 is the "background" or "ignored" class. make split works unchanged; keep_ratio: True is supported (letterbox pad is filled with ignore_index, so pad pixels don't supervise); coco_dataset: True is not supported for this task. Multi-channel inputs (RGB + thermal / depth / NIR / …) Set train.in_channels: N (default 3) to train on stacks beyond plain RGB. Supported range is N=3 (RGB) or N=4 (RGB + one extra modality, e.g. thermal, depth, NIR). Higher channel counts are not supported - cv2 / Albumentations ops cap at 4. Layout is the same; drop the stacks as .npy files (uint8 HWC arrays): data/dataset/ ├── images/ # one .npy per sample, shape (H, W, N), dtype uint8 └── labels/ # YOLO .txt (same as 3-channel case) Loader rules (see src/dl/dataset.py):
np.load is byte-faithful - channels come back exactly as you saved them. A file whose channel count doesn't match train.in_channels is skipped with a loguru.warning line (path + reason).
Mosaic re-samples another index automatically. Pretrained 3-channel backbone weights are reused: the stem conv is inflated to N input channels by tiling/averaging the RGB filters (inflate_stem_weight in src/d_fine/utils.py), so fine-tuning from COCO still works. Stem freeze (freeze_at in src/d_fine/configs.py) is auto-bypassed when in_channels > 3 so the inflated extra-channel weights can train; the size-configured freeze_at still applies for plain 3-channel RGB.
Channel-order convention: write the RGB triplet in the first three planes (channels 0..2) so they line up with the pretrained RGB stem; extra modalities go in channels 3..N-1. Example for RGB + thermal: stack as [R, G, B, T]. Example: src/etl/m3fd_to_yolo.py converts the M3FD RGB+thermal detection benchmark (PASCAL VOC XML + paired Vis//Ir/ PNGs) into this exact layout. COCO JSON format Place standard COCO JSON annotation files alongside your images folder. Splits are detected automatically by filename: data/dataset/ ├── images/ # all images ├── train.json # COCO-format annotations for train split ├── val.json # COCO-format annotations for val split └── test.json # (optional) COCO-format annotations for test split Enable COCO mode by setting coco_dataset: True in your config (see below). No CSV split generation step is needed - the splits are read directly from the JSON files. Configure Edit config.yaml - key settings: task: detect # detect | segment | sem_seg exp_name: my_exp # experiment name (used in output paths) model_name: s # n / s / m / l / x
train: root: /path/to/project # project root, will be used for outputs data_path: /path/to/dataset # folder with images/ and
labels/ (YOLO) or *.json files (COCO) coco_dataset: False # set True to use COCO JSON annotations (train.json / val.json / test.json) label_to_name: 0: class_a 1: class_b epochs: 75 batch_size: 8 img_size: [640, 640] # (h, w) Usage make split # create train/val CSV splits (test split if configured) make train # train the model make export # export to ONNX, TensorRT, OpenVINO, CoreML, LiteRT make bench # benchmark all exported models on the val set
make infer # run on test folder, save visualizations + YOLO txt predictions make check_errors # compare predictions against GT, save only mismatches (FP/FN) make test_batching # find optimal batch size for your GPU
make ov_int8 # INT8 accuracy-aware quantization for OpenVINO (can take hours) Notes:
YOLO format: make train requires train.csv and val.csv in train.data_path (generated by make split). COCO format: set coco_dataset: True - train.json and val.json are loaded directly; make split is not needed. make infer runs Torch inference on train.path_to_test_data and writes to train.infer_path.
Or run in sequence: make # train -> export -> bench (does not run split) Or run overwriting configs from CLI uv run python -m src.dl.train exp_name=my_exp Enable DDP (multi-GPU) by setting train.ddp.enabled: True and train.ddp.n_gpus: N in config. Then just run make train - it auto-launches with torchrun.