no code implementations • 14 Mar 2024 • Chris Kelly, Luhui Hu, Jiayin Hu, Yu Tian, Deshun Yang, Bang Yang, Cindy Yang, Zihao Li, Zaoshan Huang, Yuexian Zou
It seamlessly integrates various SOTA vision models and brings the automation in the selection of SOTA vision models, identifies the suitable 3D mesh creation algorithms corresponding to 2D depth maps analysis, generates optimal results based on diverse multimodal inputs such as text prompts.
no code implementations • 14 Mar 2024 • Chris Kelly, Luhui Hu, Bang Yang, Yu Tian, Deshun Yang, Cindy Yang, Zaoshan Huang, Zihao Li, Jiayin Hu, Yuexian Zou
With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achieve open-world visual perception remains an open question.
no code implementations • 10 Mar 2024 • Deshun Yang, Luhui Hu, Yu Tian, Zihao Li, Chris Kelly, Bang Yang, Cindy Yang, Yuexian Zou
Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content.
1 code implementation • 16 Nov 2023 • Chris Kelly, Luhui Hu, Cindy Yang, Yu Tian, Deshun Yang, Bang Yang, Zaoshan Huang, Zihao Li, Yuexian Zou
In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains.
no code implementations • 14 May 2020 • Chaoya Jiang, Deshun Yang, Xiaoou Chen
One part is a network for learn- ing the deep sequence representation of music tracks, and the other is a similarity estimation network which takes as input the cross- similarity matrices calculated from the deep sequences of a pair of tracks.
Ranked #1 on Cover song identification on YouTube350
2 code implementations • arXiv 2019 • Zhesong Yu, Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
We first train the network through classification strategies; the network is then used to extract music representation for cover song identification.
Ranked #3 on Cover song identification on YouTube350
1 code implementation • 20 Oct 2019 • Zehao Wang, Jingru Li, Xiaoou Chen, Zijin Li, Shicheng Zhang, Baoqiang Han, Deshun Yang
The effectiveness of the proposed framework is tested on a new dataset, its categorization of techniques is similar to our training dataset.