Pipeline: Audio Embedding using CLMR

Authors: Jael Gu

Overview

The pipeline uses a pre-trained CLMR model to extract embeddings of a given audio. It first transforms the input audio to a wave file with sample rate of 22050. Then the model splits the audio data into shorter clips with a fixed length. Finally it generates vectors of each clip, which composes the fingerprint of the input audio.

Interface

Input Arguments:

filepath:
- the input audio in .wav (audio length > 3 seconds)
- supported types: str (path to the audio)

Pipeline Output:

The Operator returns a tuple Tuple[('embs', numpy.ndarray)] containing following fields:

embs:
- embeddings of input audio
- data type: numpy.ndarray
- shape: (num_clips,512)

How to use

Install Towhee

$ pip3 install towhee

You can refer to Getting Started with Towhee for more details. If you have any questions, you can submit an issue to the towhee repository.

Install ffmpeg

$ brew install ffmpeg # for Mac

$ apt install ffmpeg # for Ubuntu

Run it with Towhee

>>> from towhee import pipeline

>>> embedding_pipeline = pipeline('towhee/audio-embedding-clmr')
>>> embedding = embedding_pipeline('path/to/your/audio')

How it works

This pipeline includes a main operator: audio-embedding (implemented as towhee/clmr-magnatagatune). The audio embedding operator encodes audio file and finally output a set of vectors of the given audio.

shiyu22 de2b0cf173 Add audio_embedding_clmr.py Signed-off-by: shiyu22 <shiyu.chen@zilliz.com>			8 Commits
README.md	1.7 KiB	Update README	5 years ago
audio_embedding_clmr.py	288 B	Add audio_embedding_clmr.py	3 years ago
audio_embedding_clmr.yaml	1.4 KiB	Add yaml	5 years ago