Transfer Learning and Optimization for Instance Segmentation Model - YOLACT

This example demonstrates a complete workflow in Kenning - from transfer learning of YOLACT (a model that detects object locations and draws their outlines/masks) for office objects, model optimization (quantization and compilation using TVM), to deployment on a GPU platform running 3 models simultaneously in real time:

The ROS2 GUI Node project will be used to demonstrate its operation.

Dependencies

To run this scenario, you will need:

  1. Hardware:

    • A camera for streaming frames

    • A CUDA-enabled NVIDIA GPU for inference acceleration

  2. Software:

    • Docker - to use a prepared environment (for ROS2)

    • UV - to quickly install Python dependencies and manage virtual environments

    • repo tool - to clone all necessary repositories

    • nvidia-container-toolkit - to provide access to the GPU in the Docker container

Info

This example requires a CUDA-enabled GPU.

Depending on the example, you may also need:

  • CUDA Toolkit (13.0 recommended), including NVCC and NVRTC, to compile CUDA code.

  • cuDNN (9 recommended) for operations that use cuDNN.

  • NVML to generate GPU reports.

Check the example to see which of these components are needed.

Installation

For transfer learning and model optimization, we can directly clone the Kenning repository from:

git clone https://github.com/antmicro/kenning.git

and perform the next steps (2, 3, 4) inside it. However, since we ultimately want to run ros2-gui-demo, we will use the repo tool for this, which will fetch other necessary repositories and dependencies alongside Kenning (including GUI Node, scripts for building the ROS 2 image, etc.):

mkdir kenning-transfer-learning && cd kenning-transfer-learning
repo init -u https://github.com/antmicro/ros2-gui-node.git -m examples/kenning-instance-segmentation/manifest.xml
repo sync -j`nproc`
mkdir build

Note

Before executing repo command you may need to set up git credential by typing into terminal:

git config --global user.email "<e-mail address>"
git config --global user.name "Name Surname"

Configuration Files

Everything in Kenning revolves around configuration files, which store configurations for various scenarios, such as model optimization. Let’s enter the directory where Kenning is located (kenning/). In the scripts/configs/ directory, we can find the pytorch-open-images-yolact-tvm-gpu.yaml file, which contains the YOLACT model configuration for report generation, optimization, training, and benchmarking. The individual sections are taken into account when invoking the relevant scenarios using the kenning [scenario] [params] command, e.g., kenning train --cfg path/file.yaml. In this case, only model_wrapper and dataset will be taken into account:

model_wrapper:
  type: YOLACTOpenImages
  parameters:
    # Output of fine-tuning; also used as input for the optimize stage
    model_path: ./build/yolact_finetuned.pth
    # Starting checkpoint (backbone frozen, only heads are trained)
    pretrained_weights_path: kenning:///models/instance_segmentation/yolact_resnet50.pth
    freeze_backbone: true
    learning_rate: 0.0004
    batch_size: 32
    num_epochs: 1
    logdir: ./build/logs

dataset:
  type: OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    classes: coco
    download_annotations_type: train
    image_width: 550
    image_height: 550

More information about this can be found here: Defining optimization pipelines in Kenning. Unfortunately, there is currently no single place describing all block arguments, so you need to search for them manually in the corresponding class definition file.

Transfer Learning

Transfer learning involves loading a pretrained model (weights) that was trained for task A and, after minor modifications, adapting it so that it can be used for task B. In practice, this means usually adding or modifying the prediction head (e.g., changing the number of recognized classes), freezing the rest of the model so it isn’t trained, and then training this model for a few epochs. The next step is reducing the learning rate to avoid damaging the previously learned weights, unfreezing them, and training for a few epochs so the model adapts to the new head.

Selecting Classes for Recognition

The YOLACT model was originally trained on the COCO dataset (part of the architecture, namely ResNet-50, was previously trained on ImageNet). It is therefore natural that we want to select objects (and a dataset) that roughly resemble those from the original training - general real-world objects, rather than, for example, computer games.

For this purpose, we will use the OpenImages dataset implemented in Kenning and select a subset of classes from it - objects that can be found in an office.

Note

OpenImages dataset labels are encoded (they do not contain direct names like Tree), and the file with all classes and their encodings can be found here: OpenImages Class Descriptions.

These classes will be:

/m/02jvh9,Mug
/m/02p0tk3,Human body
/m/0k1tl,Pen
/m/020lf,Computer mouse
/m/01m2v,Computer keyboard
/m/02522,Computer monitor
/m/050k8,Mobile phone
/m/04dr76w,Bottle
/m/02dl1y,Hat
/m/0242l,Coin
/m/080hkjn,Handbag

Let’s save these classes to the transfer_learning.csv file, preferably in the Kenning root directory (kenning/).

Adapting the Prediction Head

Now it’s time to fine-tune YOLACT so that it recognizes new objects. Let’s open the pytorch-open-images-yolact-tvm-gpu.yaml file and modify the dataset section - let’s change classes: coco to classes: "transfer_learning.csv":

dataset:
  type: OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    classes: coco
    download_annotations_type: train
    image_width: 550
    image_height: 550

Let’s change num_epochs to 10. This will train the head for 10 epochs:

model_wrapper:
  type: YOLACTOpenImages
  parameters:
    # Output of fine-tuning; also used as input for the optimize stage
    model_path: ./build/yolact_finetuned.pth
    # Starting checkpoint (backbone frozen, only heads are trained)
    pretrained_weights_path: kenning:///models/instance_segmentation/yolact_resnet50.pth
    freeze_backbone: true
    learning_rate: 0.0004
    batch_size: 32
    num_epochs: 1
    logdir: ./build/logs

Next, let’s run the container with installed dependencies, including ROS2 and TVM (for simplicity, to avoid creating a virtual environment, since we will be using this container later anyway), using:

# Building the image, this may take a while
./src/gui_node/environments/build-docker.sh gpu

# Allow non-network local connections to X11 so that the GUI can be started from the Docker container
xhost +local:

# Run the container
./src/gui_node/environments/run-docker.sh gpu

and let’s install Kenning and the required libraries with:

uv pip install --project ./kenning -e "./kenning[object_detection,pose_estimation,onnxruntime_gpu,torch]" --group tvm-cuda

or when not using ROS2, create a virtual environment:

uv venv
source .venv/bin/activate

Note

See the Kenning installation section for information about the Python versions currently supported.

and run:

uv pip install --project ./kenning -e "./kenning[object_detection, torch]" --group tvm-cuda

Then, let’s run the training process for the head itself (it will be modified automatically based on the number of dataset classes):

kenning train --cfg kenning/scripts/configs/pytorch-open-images-yolact-tvm-gpu.yaml

Depending on the available hardware, it may take a while. Our model will be saved to ./build/yolact_finetuned.pth.

Fine-tuning the Rest of the Model

The next step is to unfreeze the remaining part of the model (changing freeze_backbone: false), increase the number of epochs (num_epochs: 40), and decrease the learning rate (learning_rate: 0.00005). Additionally, the recently trained model (pretrained_weights_path: ./build/yolact_finetuned.pth) should be used as a starting point:

model_wrapper:
  type: YOLACTOpenImages
  parameters:
    # Output of fine-tuning; also used as input for the optimize stage
    model_path: ./build/yolact_finetuned.pth
    # Starting checkpoint (backbone frozen, only heads are trained)
    pretrained_weights_path: kenning:///models/instance_segmentation/yolact_resnet50.pth
    freeze_backbone: true
    learning_rate: 0.0004
    batch_size: 32
    num_epochs: 1
    logdir: ./build/logs

Let’s train the model again, this time in its entirety:

kenning train --cfg kenning/scripts/configs/pytorch-open-images-yolact-tvm-gpu.yaml

The resulting model will be located in ./build/final_yolact.pth.

Optimization of the Trained Model

We can speed up our model almost 5-fold by applying int8 quantization and compiling with TVM (using appropriate flags):

  • Quantization consists of reducing the precision of the model’s weights and activations (e.g., from FP32 floating-point format to 8-bit integer format INT8). This translates into faster matrix operations, lower RAM/VRAM usage, and reduced energy consumption. However, quantization is a lossy optimization, meaning that a slight drop in model accuracy should be expected.

  • TVM compilation involves converting the network into an Intermediate Representation (IR), where computational graph optimizations are performed, including operation fusion (combining consecutive layers into one, which reduces memory transfers). The compiler then generates machine code optimized for a specific hardware architecture (e.g., x86, ARM, RISC-V, CUDA), utilizing its specific instructions (e.g., AVX-512, NEON, or Tensor Cores) to maximize inference performance.

In Kenning, we can do this incredibly easily. Simply change target: cuda to target: cuda -arch=sm_86, where sm_86 specifies the architecture (Compute Capability) of our graphics card - in our case, the RTX 3090. You can check your GPU’s architecture using the following command:

nvidia-smi --query-gpu=compute_cap --format=csv

Note

Convert the returned value (e.g., 8.6) into the sm_XX format by removing the decimal point and adding the sm_ prefix (for 8.6, this becomes sm_86).

Quantization (in this example, Post-Training Quantization) is enabled using use_int8_precision: true. We apply PTQ with calibration, where dataset_percentage: 0.001 feeds a small slice of the dataset to the compiler. This sample is required to collect activation statistics (min/max value ranges) and calculate accurate quantization scaling factors. The final optimizers section for optimization is:

optimizers:
  - type: ONNXCompiler
    parameters:
      compiled_model_path: ./build/yolact.onnx

  - type: TVMCompiler
    parameters:
      model_framework: onnx
      # This can speed things up by almost 4x.
      target: cuda -arch=sm_86
      opt_level: 3
      use_int8_precision: true
      # Small calibration slice to keep quantization fast and avoid out-of-memory error
      dataset_percentage: 0.001
      compiled_model_path: ./build/yolact_tvm_int8.tar

Now, let’s run the optimization scenario using the command:

kenning optimize --cfg kenning/scripts/configs/pytorch-open-images-yolact-tvm-gpu.yaml

Our optimized model is located in ./build/yolact_tvm_int8.tar.

Real-Time Usage Example

ROS2 GUI Node is a project created for visualizing data from ROS 2. ROS 2 itself is a robotics middleware based on a publish-subscribe (pub/sub) pattern and a node-based architecture. Conceptually, it functions much like a microservices framework, enabling the development of efficient and modular applications for edge devices. In this architecture, every module - from the camera, through individual AI models, to the graphical user interface (GUI) - runs as a separate, independent node.

Let’s replace the entire contents of the kenning-instance-segmentation.yaml, located in the src/gui_node/examples/kenning-multimodel-demo directory, with the following configuration:

- type: kenning.dataproviders.ros2_camera_node_data_provider.ROS2CameraNodeDataProvider
  parameters:
    topic_name: camera_frame
    output_memory_layout: NCHW
    output_width: 550
    output_height: 550
  outputs:
    frame: cam_frame
    frame_original: cam_frame_original

- type: kenning.runners.modelruntime_runner.ModelRuntimeRunner
  parameters:
    model_wrapper:
      type: kenning.modelwrappers.instance_segmentation.pytorch_yolact.YOLACTOpenImages
      parameters:
        model_path: build/final_yolact.pth
        max_detections: 100
        score_threshold: 0.2
    runtime:
      type: kenning.runtimes.tvm.TVMRuntime
      parameters:
        save_model_path: build/yolact_tvm_int8.tar
        target_device_context: cuda
  inputs:
    input: cam_frame
  outputs:
    segmentation_output: predictions

- type: kenning.outputcollectors.ros2_yolact_outputcollector.ROS2YolactOutputCollector
  parameters:
    topic_name: instance_segmentation_kenning
    input_color_format: BGR
    input_memory_layout: NCHW
  inputs:
    frame_original: cam_frame_original
    output: predictions

In short, we changed the runtime to TVM, updated the model wrapper to point to the appropriate class, modified file paths (including the one for the optimized model), and added a dataset section from which the model will retrieve class names and map them to the model’s output (integers).

Let’s build the nodes:

source /opt/ros/$ROS_DISTRO/setup.bash

colcon build --base-paths src --cmake-args -DBUILD_KENNING_MULTIMODEL_DEMO=y -DPython3_EXECUTABLE=/opt/venv/bin/python3

And finally, let’s run the demo:

source install/setup.sh

ros2 launch gui_node kenning-multimodel-demo.py use_gui:=True

Summary

In this example, we used Kenning for YOLACT transfer learning and squeezed as much performance out of it as possible. Finally, we ran it alongside MMPose and DINOv2. Thanks to Kenning, optimization and transfer learning are incredibly simple and convenient, without any loss in performance.


Last update: 2026-10-01