Microsoft ends support for Internet Explorer on June 16, 2022.
We recommend using one of the browsers listed below.

  • Microsoft Edge(Latest version) 
  • Mozilla Firefox(Latest version) 
  • Google Chrome(Latest version) 
  • Apple Safari(Latest version) 

Please contact your browser provider for download and installation instructions.

Open search panel Close search panel Open menu Close menu

September 14, 2026

Information

3 papers from NTT accepted to ICIP, a leading international conference in the field of image processing, computer vision imaging

NTT accepted 3 papers to ICIP 2026(International Conference on Image Processing 2026), the world's largest international flagship conference in the fields of image processing, image recognition (vision) and imaging, held in Tampere, Finland from September 13 to 17th. Organized by the IEEE Signal Processing Society, an international society in electrical and electronic engineering, this conference has been held every year since 1994, and 2026 will be the 33rd. The selected papers are as follows.

Abbreviated names of the laboratories:
HI: Human Informatics Labs., NTT, Inc.
CD: Computer and Data Science Labs., NTT, Inc.

■End-to-End Learning of Metalens-Based Compressive Sensing and Differentiable Coding for Hyperspectral Image Transmission

Takayuki Sasaki (CD), Yoko Sogabe (CD), Kazuya Hayase (CD), Masaki Kitahara (CD), Yukihiro Bandoh (Shimonoseki City Univ.)

Hyperspectral Imaging (HSI) captures rich spectral information beyond the capabilities of the human eye. By utilizing remote sensing (RS) from drones and satellites, HSI is expected to be deployed across a wide range of fields, including environmental monitoring and smart agriculture. In such RS applications, captured data must be transmitted to the ground under high-compression and low-computation constraints, making re-compression via an image codec essential. However, conventional, naive re-compression amplifies coding artifacts during the reconstruction process, which degrades HSI quality and leads to analysis failures in downstream application phases. In this report, we present a newly established technology optimized to suppress these coding artifacts by connecting a differentiable image codec and HSI reconstructor end-to-end for joint training. Compared to conventional re-compression methods, this approach successfully reduces the transmitted data size by approximately 40% while preserving high image quality.

■Clip-Level Uncertainty and Temporal-Aware Active Learning for End-to-End Multi-Object Tracking

Riku Inoue (HI), Shogo Sato (HI), Kazuhiko Murasaki (HI), Tomoyasu Shimada (HI), Toshihiko Nishimura (HI), Ryuichi Tanida (HI)

We propose a method for efficiently selecting training samples that contribute to improving the accuracy of multi-object tracking in video. Multi-object tracking is a fundamental technology for continuously tracking people, vehicles, and other objects in video footage. However, training a high-accuracy tracker typically requires dense manual annotations, including the position and identity of each object across video frames. This creates a significant cost barrier when adapting tracking technology to new object categories or deployment environments. Our proposed method evaluates the importance of training samples using short video clips. It quantitatively assesses factors such as variations in tracking confidence and the consistency between forward and backward tracking results. By prioritizing highly informative samples while maintaining temporal diversity, the method constructs an effective training sample set with a reduced amount of annotated data.In evaluation experiments, we confirmed that, on some datasets, using only 50% of the training data achieved tracking accuracy comparable to using the full dataset.

■Social Group Activity Recognition from Still Images Using Conditional Token Sequence Generation

Shota Orihashi (HI), Taiga Yamane (HI), Naoki Makishima (HI), Mana Ihori (HI), Satoshi Suzuki (HI), Tomohiro Tanaka (HI), Ryo Masumura (HI)

Social group activity recognition is a task that detects humans and identifies multiple human groups and their activities in a scene. Existing methods rely on video input for this task. We extend social group activity recognition from video to still images to enable broader application. Since still images lack motion information, detecting humans and their interactions is challenging. As a result, reliable recognition usually requires large annotated datasets, which are not easy to obtain. To address this issue, our key idea is to train a model simultaneously on tasks with rich still-image datasets that localize objects, such as human detection and object detection, to strengthen the model's ability to find and localize objects, including humans. To this end, we introduce a multitask conditional token sequence generation model that, given a token specifying the target task, outputs the task result as a token sequence. To represent social group activity recognition results, we propose an activity label token that simultaneously represents the activity label and the group boundary. The proposed method allows the model to be trained on multiple object localization tasks and improves performance. Experiments on public benchmark datasets demonstrate the effectiveness of our approach.

Information is current as of the date of issue of the individual topics.
Please be advised that information may be outdated after that point.