Filter
SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View Projection
- Computer Vision
- CVPR
- 2025
GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Tracking
- Multi-Object Tracking
- CVPR
- 2025
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
- Speech Processing
- ICASSP
- 2024
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
- Computer Vision
- CVPR
- 2024
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
- Keyword Spotting
- IEEE Signal Processing Letters
- 2024
Who Should Have Been Focused: Transferring Attention-Based Knowledge from Future Observations for Trajectory Prediction
- Trajectory Prediction
- ICPR
- 2024
Forbes: Face Obfuscation Rendering via Backpropagation Refinement Scheme
- Face obfuscation
- ECCV
- 2024
Self-training ASR Guided by Unsupervised ASR Teacher
- Speech Recognition
- Interspeech
- 2024
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
- Natural Language Processing
- COLING
- 2024
Joint Appearance and Motion Model with Temporal Transformer for Multiple Object Tracking
- Computer Vision
- IEEE Access
- 2023
Boosting Unknown-number Speaker Separation With Transformer Decoder-based Attractor
- Speech Separation
- ICASSP
- 2024
Voxtlm: Unified Decoder-only Models for Consolidating Speech Recognition/Synthesis and Speech/Text Continuation Tasks
- Speech Recognition, Speech Synthesis
- ICASSP
- 2024
Learning Contextualized Representation On Discrete Space Via Hierarchical Product Quantization
- Speech Recognition
- ICASSP
- 2024
TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech Separation
- Speech Separation
- IEEE/ACM TASLP
- 2023
That's What Said: Fully-Controllable Talking Face Generation
- Computer Vision, Pattern Recognition
- ACM/MM
- 2023
Luminance-aware Color Transform for Multiple Exposure Correction
- Computer Vision, Ehancement
- ICCV
- 2023
SlaBins: Fisheye Depth Estimation using Slanted Bins on Road Environments
- Computer Vision, 3D
- ICCV
- 2023
SpeedFormer: Learning Speed Profiles with Upper and Lower Boundary Constraints Based on Transformer
- Motion Planning
- IROS
- 2023
Factspeech: Speaking a Foreign Language Pronunciation Using Only Your Native Characters
- Speech Synthesis
- Interspeech
- 2023
MiLO: Multi-task Learning with Localization Ambiguity Suppression for Occupancy Prediction
- Computer Vision, 3D
- CVPRW
- 2023