Optimizing and comparing an instance segmentation model

This example demonstrates how to optimize an instance segmentation model and compare different optimizers on CPU and GPU.

Base model used for demonstration is going to be YOLACT, short for “You Only Look At CoefficientTs” which is a fully convolutional model for real-time instance segmentation. We will also need a dataset for evaluation purposes - in this case it is going to be OpenImagesDatasetV6.

The model will first be optimized for a given target using Kenning optimizers (TVM, ONNX), and finally compiled into a deployable, runtime-specific artifact. Model will be deployed on CPU and GPU using following Kenning compilers:

Dependencies

To run this scenario, you will need:

  1. Hardware:

    • A CUDA-enabled NVIDIA GPU

  2. Software:

    • UV - to quickly install Python dependencies and manage virtual environments

Experiments on GPU

In this section, the model will first be optimized for GPU execution using TVM and ONNX Runtime, and then benchmarked to compare the resulting performance.

Info

This example requires a CUDA-enabled GPU.

Depending on the example, you may also need:

  • CUDA Toolkit (13.0 recommended), including NVCC and NVRTC, to compile CUDA code.

  • cuDNN (9 recommended) for operations that use cuDNN.

  • NVML to generate GPU reports.

Check the example to see which of these components are needed.

Installation

First, let’s install the dependencies needed to run this scenario on GPU, and we’ll get to CPU afterwards.

Let’s install Kenning:

git clone https://github.com/antmicro/kenning.git

and the required libraries:

uv pip install \
  --project ./kenning \
  --group tvm-cuda \
  "./kenning[object_detection,onnxruntime_gpu,reports]"

(GPU) Optimization

Optimization is the process of converting the trained model into a form better suited for fast, efficient inference on a given target. With TVMCompiler, the model is compiled through TVM’s compiler stack into low-level, target-specific code (e.g. CUDA kernels for GPU, LLVM-generated code for CPU), applying optimizations such as operator fusion and memory layout tuning. With ONNXCompiler, the model is instead converted into an optimized ONNX graph, which ONNXRuntime can then execute directly using its own set of hardware-specific execution providers (e.g. CUDAExecutionProvider, CPUExecutionProvider).

(GPU) TVM Optimization

To optimize the model on GPU using TVM, the yolact-tvm-gpu-detection.yml file has been prepared, which looks as follows:

Listing 15 yolact-tvm-gpu-detection.yml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
model_wrapper:
  type: kenning.modelwrappers.instance_segmentation.yolact.YOLACT
  parameters:
    model_path: kenning:///models/instance_segmentation/yolact.onnx
dataset:
  type: kenning.datasets.open_images_dataset.OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    image_width: 550
    image_height: 550
optimizers:
- type: kenning.optimizers.tvm.TVMCompiler
  parameters:
    model_framework: onnx
    target: "cuda -libs=cudnn,cublas"
    opt_level: 3
    compile_use_vm: false
    output_conversion_function: default
    target_host: null
    compiled_model_path: ./build/compiled-model.tar
runtime:
  type: kenning.runtimes.tvm.TVMRuntime
  parameters:
    save_model_path: ./build/compiled-model.tar
    target_device_context: cuda
    runtime_use_vm: false

In this scenario:

  • model_path points to a location of the YOLACT model in ONNX format. It can be either local file or a remote URL. kenning:// is a special schema for Kenning’s demonstration models.

  • dataset tells to use Open Images dataset. The model will be downloaded to ./build/OpenImagesDatasetV6. The task field allows to specify whether the dataset is used for instance segmentation or object detection.

  • optimizers contains only one element - TVMCompiler. In there we specify input model framework (onnx), and tell to use cuda target with opt_level equal to 3 (applying all possible optimizations not directly affecting model’s output).

  • runtime tells Kenning to use TVMRuntime for model execution, on GPU target.

To learn more about config files see Optimization Pipelines.

Setting target to cuda -libs=cudnn,cublas tells TVMCompiler to compile the model for CUDA execution, using cuDNN and cuBLAS kernels for the supported operations. Setting target_device_context to cuda tells TVMRuntime to run inference on the GPU instead of the CPU.

To run the optimization and test the model, execute:

kenning optimize test \
    --cfg yolact-tvm-gpu-detection.yml \
    --measurements ./build/yolact-gpu-tvm.json \
    --verbosity INFO

(GPU) ONNX Optimization

Now it’s time for ONNX. Compared to TVM, the difference is that we change the optimizer to ONNXCompiler and the runtime to ONNXRuntime. The yolact-onnx-gpu-detection.yml file has been prepared for this:

Listing 16 yolact-onnx-gpu-detection.yml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
platform:
  type: CUDAPlatform
  parameters:
    name: jetson_thor_devkit
model_wrapper:
  type: kenning.modelwrappers.instance_segmentation.yolact.YOLACT
  parameters:
    model_path: kenning:///models/instance_segmentation/yolact.onnx
dataset:
  type: kenning.datasets.open_images_dataset.OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    image_width: 550
    image_height: 550
optimizers:
- type: kenning.optimizers.onnx.ONNXCompiler
  parameters:
      compiled_model_path: ./build/compiled_model.onnx
runtime:
  type: kenning.runtimes.onnx.ONNXRuntime
  parameters:
    save_model_path: ./build/compiled_model.onnx
    execution_providers:
    - CUDAExecutionProvider
    - CPUExecutionProvider
report:
  type: MarkdownReport
  parameters:
    measurements: ./build/yolact-onnx-gpu-detection-cfg.json
    report_path: ./build/yolact-onnx-gpu-detection/report.md
    report_name: "Sample ONNX model on GPU inference report"
    report_types:
      - performance
      - classification
    to_html: true

execution_providers in ONNXRuntime holds a list of possible layer executors, starting from the most preferred one. For GPU execution, CUDAExecutionProvider should be listed first, with CPUExecutionProvider as a fallback for operations without CUDA support.

Run the scenario as follows:

kenning optimize test \
    --cfg yolact-onnx-gpu-detection.yml \
    --measurements ./build/yolact-gpu-onnx.json \
    --verbosity INFO

Comparison of GPU runtimes

To create a comparison report comparing performance and model quality for the above optimizers, run:

kenning report \
    --report-path build/yolact-report/report.md \
    --report-name "YOLACT detection report" \
    --root-dir build/yolact-report \
    --img-dir build/yolact-report/imgs \
    --report-types performance detection \
    --measurements build/yolact-gpu-tvm.json build/yolact-gpu-onnx.json
Mean performance comparison

Figure 6 Model size, speed and quality comparison for two YOLACT Optimizers running on CUDA GPU

Experiments on CPU

Most of the steps for CPU will look nearly identical to those for GPU - the CPU section that follows mainly highlights the small modifications needed to switch from GPU to CPU execution.

Installation

Since we currently have TVM and ONNX Runtime packages installed for GPU, we need to uninstall them and install the CPU-compatible versions instead. We can install CPU-only replacements with:

uv pip install "./kenning[object_detection, tvm, onnxruntime, reports]"

(CPU) Optimization

(CPU) TVM Optimization

The yolact-tvm-cpu-detection.yml file has been prepared to optimize the model with TVM on CPU, and looks as follows:

Listing 17 yolact-tvm-cpu-detection.yml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
model_wrapper:
  type: kenning.modelwrappers.instance_segmentation.yolact.YOLACT
  parameters:
    model_path: kenning:///models/instance_segmentation/yolact.onnx
dataset:
  type: kenning.datasets.open_images_dataset.OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    image_width: 550
    image_height: 550
optimizers:
- type: kenning.optimizers.tvm.TVMCompiler
  parameters:
    model_framework: onnx
    target: llvm
    opt_level: 3
    compile_use_vm: false
    compiled_model_path: ./build/compiled-model.tar
runtime:
  type: kenning.runtimes.tvm.TVMRuntime
  parameters:
    save_model_path: ./build/compiled-model.tar
    target_device_context: cpu
    runtime_use_vm: false

Compared to the GPU scenario, the only differences are the target and target_device_context - here set to llvm and cpu respectively, instead of cuda -libs=cudnn,cublas and cuda.

To optimize and test the defined scenario, run:

kenning optimize test \
    --cfg yolact-tvm-cpu-detection.yml \
    --measurements ./build/yolact-tvm.json \
    --verbosity INFO

(CPU) ONNX Optimization

Similarly to the GPU case, switching to a different runtime is a matter of changing several lines in the scenario, as shown below:

Listing 18 yolact-onnx-cpu-detection.yml
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
platform:
  type: LocalPlatform
model_wrapper:
  type: kenning.modelwrappers.instance_segmentation.yolact.YOLACT
  parameters:
    model_path: kenning:///models/instance_segmentation/yolact.onnx
dataset:
  type: kenning.datasets.open_images_dataset.OpenImagesDatasetV6
  parameters:
    dataset_root: ./build/OpenImagesDatasetV6
    inference_batch_size: 1
    task: instance_segmentation
    image_width: 550
    image_height: 550
optimizers:
- type: kenning.optimizers.onnx.ONNXCompiler
  parameters:
    compiled_model_path: ./build/compiled_model.onnx
runtime:
  type: kenning.runtimes.onnx.ONNXRuntime
  parameters:
    save_model_path: ./build/compiled_model.onnx
    execution_providers:
    - CPUExecutionProvider

This scenario can be executed with:

kenning optimize test \
    --cfg yolact-onnx-cpu-detection.yml \
    --measurements ./build/yolact-onnx.json \
    --verbosity INFO

Comparison of CPU runtimes

To create a comparison report comparing performance and model quality for the above optimizers, run:

kenning report \
    --report-path build/yolact-report/report.md \
    --report-name "YOLACT detection report" \
    --root-dir build/yolact-report \
    --img-dir build/yolact-report/imgs \
    --report-types performance detection \
    --measurements build/yolact-tvm.json build/yolact-onnx.json
Mean performance comparison

Figure 7 Sample comparison plot demonstrating model size, speed and quality for two YOLACT Optimizers

Summary

In this example, we optimized and deployed the YOLACT instance segmentation model on both GPU and CPU, using two different Kenning optimizers - TVMCompiler and ONNXCompiler - paired with their corresponding runtimes, TVMRuntime and ONNXRuntime.

The overall workflow was the same across both targets and both optimizers: prepare a scenario configuration, run kenning optimize test to optimize and benchmark the model, and generate a comparison report with kenning report. Switching between GPU and CPU, or between TVM and ONNX, required only small, targeted changes to the scenario - primarily the target/target_device_context fields for TVM, and the execution_providers list for ONNX Runtime.

The resulting reports allow for a direct comparison of model size, inference speed and detection quality across all four configurations (TVM/ONNX × GPU/CPU), making it easier to choose the best-performing setup for a given deployment target.


Last update: 2026-10-01