Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is a schematic diagram of a structure of a GAN- based speech synthesis model according to an embodiment of the present disclosure;
FIG. 2
FIG. 2 is a schematic diagram of an operation flow of a GAN-based speech synthesis model according to an embodiment of the present disclosure;
FIG. 3
FIG. 3 is a flowchart of a speech synthesis method implemented by a speech synthesis model according to an embodiment of the present disclosure; and
FIG. 4
FIG. 4 is a flowchart of a training method for a GAN- based speech synthesis model according to an embodiment of the present disclosure.
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
1 independent · 6 dependent
1
IndependentGAN-based speech synthesis model
A GAN-based speech synthesis model, comprising a generator, configured to be obtained by being trained based on a first discrimination loss for indicating a discrimination loss of the generator and a second discrimination loss for indicating a mean square error between the generator and a preset discriminator; and a vocoder, configured to synthesize target audio corre-sponding to-be-converted text from a target Mel-frequency spectrum, wherein the generator comprises: a feature encoding layer, configured to obtain a text feature based on a text vector, the text vector being obtained by processing the to-be-converted text; an attention mechanism layer, configured to calculate, based on a sequence order of the text feature, a rel-evance between the text feature at a current position and an audio feature within a preset range, and deter-mine contribution values of each text feature relative to different audio features within the preset range, the audio feature being used for indicating an audio feature corresponding to a pronunciation object preset by the generator; and B₁ a feature decoding layer, configured to match the audio feature corresponding to the text feature based on the contribution value, and output the target Mel-frequency spectrum by the audio feature.
2
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein the generator adopts a self-cycle structure or a non-self-cycle structure.
3
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein for implementing a speech synthesis method, the model is configured to: acquire the to-be-converted text; convert the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitize the text phoneme to obtain text data; convert the text data into a text vector; and process the text vector into the target audio corresponding to the to-be-converted text.
7
Dependent← claim 1GAN-based speech synthesis model
A GAN-based speech synthesis method, applicable to the speech synthesis model according to claim 1, compris-ing: acquiring to-be-converted text; converting the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitizing the text phoneme to obtain text data; converting the text data into a text vector; and inputting the text vector into the speech synthesis model to obtain target audio corresponding to the to-be-converted text. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
GAN-based speech synthesis model
vocodervocoder
generatorgenerator
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 14
US 2021/0312243 A12021/0312243 A1 * 10/2021 Wang...................... G06T 7/194examiner
US 2022/0208355 A12022/0208355 A1 * 6/2022 Li......................... G06T 7/0016examiner
US 2022/0392428 A12022/0392428 A1 * 12/2022 Fernandez Guajardo...................examiner
CN 111627418 ACN 111627418 A 9/2020
Patent
Atlas literature
Patent
US 11,817,079 B1
GAN-BASED SPEECH SYNTHESIS MODEL AND TRAINING METHOD
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is a schematic diagram of a structure of a GAN- based speech synthesis model according to an embodiment of the present disclosure;
FIG. 2
FIG. 2 is a schematic diagram of an operation flow of a GAN-based speech synthesis model according to an embodiment of the present disclosure;
FIG. 3
FIG. 3 is a flowchart of a speech synthesis method implemented by a speech synthesis model according to an embodiment of the present disclosure; and
FIG. 4
FIG. 4 is a flowchart of a training method for a GAN- based speech synthesis model according to an embodiment of the present disclosure.
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
1 independent · 6 dependent
1
IndependentGAN-based speech synthesis model
A GAN-based speech synthesis model, comprising a generator, configured to be obtained by being trained based on a first discrimination loss for indicating a discrimination loss of the generator and a second discrimination loss for indicating a mean square error between the generator and a preset discriminator; and a vocoder, configured to synthesize target audio corre-sponding to-be-converted text from a target Mel-frequency spectrum, wherein the generator comprises: a feature encoding layer, configured to obtain a text feature based on a text vector, the text vector being obtained by processing the to-be-converted text; an attention mechanism layer, configured to calculate, based on a sequence order of the text feature, a rel-evance between the text feature at a current position and an audio feature within a preset range, and deter-mine contribution values of each text feature relative to different audio features within the preset range, the audio feature being used for indicating an audio feature corresponding to a pronunciation object preset by the generator; and B₁ a feature decoding layer, configured to match the audio feature corresponding to the text feature based on the contribution value, and output the target Mel-frequency spectrum by the audio feature.
2
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein the generator adopts a self-cycle structure or a non-self-cycle structure.
3
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein for implementing a speech synthesis method, the model is configured to: acquire the to-be-converted text; convert the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitize the text phoneme to obtain text data; convert the text data into a text vector; and process the text vector into the target audio corresponding to the to-be-converted text.
7
Dependent← claim 1GAN-based speech synthesis model
A GAN-based speech synthesis method, applicable to the speech synthesis model according to claim 1, compris-ing: acquiring to-be-converted text; converting the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitizing the text phoneme to obtain text data; converting the text data into a text vector; and inputting the text vector into the speech synthesis model to obtain target audio corresponding to the to-be-converted text. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
GAN-based speech synthesis model
vocodervocoder
generatorgenerator
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 14
US 2021/0312243 A12021/0312243 A1 * 10/2021 Wang...................... G06T 7/194examiner
US 2022/0208355 A12022/0208355 A1 * 6/2022 Li......................... G06T 7/0016examiner
US 2022/0392428 A12022/0392428 A1 * 12/2022 Fernandez Guajardo...................examiner
CN 111627418 ACN 111627418 A 9/2020
Patent
Atlas literature
Patent
US 11,817,079 B1
GAN-BASED SPEECH SYNTHESIS MODEL AND TRAINING METHOD
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is a schematic diagram of a structure of a GAN- based speech synthesis model according to an embodiment of the present disclosure;
FIG. 2
FIG. 2 is a schematic diagram of an operation flow of a GAN-based speech synthesis model according to an embodiment of the present disclosure;
FIG. 3
FIG. 3 is a flowchart of a speech synthesis method implemented by a speech synthesis model according to an embodiment of the present disclosure; and
FIG. 4
FIG. 4 is a flowchart of a training method for a GAN- based speech synthesis model according to an embodiment of the present disclosure.
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
1 independent · 6 dependent
1
IndependentGAN-based speech synthesis model
A GAN-based speech synthesis model, comprising a generator, configured to be obtained by being trained based on a first discrimination loss for indicating a discrimination loss of the generator and a second discrimination loss for indicating a mean square error between the generator and a preset discriminator; and a vocoder, configured to synthesize target audio corre-sponding to-be-converted text from a target Mel-frequency spectrum, wherein the generator comprises: a feature encoding layer, configured to obtain a text feature based on a text vector, the text vector being obtained by processing the to-be-converted text; an attention mechanism layer, configured to calculate, based on a sequence order of the text feature, a rel-evance between the text feature at a current position and an audio feature within a preset range, and deter-mine contribution values of each text feature relative to different audio features within the preset range, the audio feature being used for indicating an audio feature corresponding to a pronunciation object preset by the generator; and B₁ a feature decoding layer, configured to match the audio feature corresponding to the text feature based on the contribution value, and output the target Mel-frequency spectrum by the audio feature.
2
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein the generator adopts a self-cycle structure or a non-self-cycle structure.
3
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein for implementing a speech synthesis method, the model is configured to: acquire the to-be-converted text; convert the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitize the text phoneme to obtain text data; convert the text data into a text vector; and process the text vector into the target audio corresponding to the to-be-converted text.
7
Dependent← claim 1GAN-based speech synthesis model
A GAN-based speech synthesis method, applicable to the speech synthesis model according to claim 1, compris-ing: acquiring to-be-converted text; converting the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitizing the text phoneme to obtain text data; converting the text data into a text vector; and inputting the text vector into the speech synthesis model to obtain target audio corresponding to the to-be-converted text. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
GAN-based speech synthesis model
vocodervocoder
generatorgenerator
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 14
US 2021/0312243 A12021/0312243 A1 * 10/2021 Wang...................... G06T 7/194examiner
US 2022/0208355 A12022/0208355 A1 * 6/2022 Li......................... G06T 7/0016examiner
US 2022/0392428 A12022/0392428 A1 * 12/2022 Fernandez Guajardo...................examiner
CN 111627418 ACN 111627418 A 9/2020
Patent
Atlas literature
Patent
US 11,817,079 B1
GAN-BASED SPEECH SYNTHESIS MODEL AND TRAINING METHOD
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is a schematic diagram of a structure of a GAN- based speech synthesis model according to an embodiment of the present disclosure;
FIG. 2
FIG. 2 is a schematic diagram of an operation flow of a GAN-based speech synthesis model according to an embodiment of the present disclosure;
FIG. 3
FIG. 3 is a flowchart of a speech synthesis method implemented by a speech synthesis model according to an embodiment of the present disclosure; and
FIG. 4
FIG. 4 is a flowchart of a training method for a GAN- based speech synthesis model according to an embodiment of the present disclosure.
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
1 independent · 6 dependent
1
IndependentGAN-based speech synthesis model
A GAN-based speech synthesis model, comprising a generator, configured to be obtained by being trained based on a first discrimination loss for indicating a discrimination loss of the generator and a second discrimination loss for indicating a mean square error between the generator and a preset discriminator; and a vocoder, configured to synthesize target audio corre-sponding to-be-converted text from a target Mel-frequency spectrum, wherein the generator comprises: a feature encoding layer, configured to obtain a text feature based on a text vector, the text vector being obtained by processing the to-be-converted text; an attention mechanism layer, configured to calculate, based on a sequence order of the text feature, a rel-evance between the text feature at a current position and an audio feature within a preset range, and deter-mine contribution values of each text feature relative to different audio features within the preset range, the audio feature being used for indicating an audio feature corresponding to a pronunciation object preset by the generator; and B₁ a feature decoding layer, configured to match the audio feature corresponding to the text feature based on the contribution value, and output the target Mel-frequency spectrum by the audio feature.
2
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein the generator adopts a self-cycle structure or a non-self-cycle structure.
3
Dependent← claim 1GAN-based speech synthesis model
The GAN-based speech synthesis model according to claim 1, wherein for implementing a speech synthesis method, the model is configured to: acquire the to-be-converted text; convert the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitize the text phoneme to obtain text data; convert the text data into a text vector; and process the text vector into the target audio corresponding to the to-be-converted text.
7
Dependent← claim 1GAN-based speech synthesis model
A GAN-based speech synthesis method, applicable to the speech synthesis model according to claim 1, compris-ing: acquiring to-be-converted text; converting the to-be-converted text into a text phoneme based on spelling of the to-be-converted text; digitizing the text phoneme to obtain text data; converting the text data into a text vector; and inputting the text vector into the speech synthesis model to obtain target audio corresponding to the to-be-converted text. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
GAN-based speech synthesis model
vocodervocoder
generatorgenerator
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 14
US 2021/0312243 A12021/0312243 A1 * 10/2021 Wang...................... G06T 7/194examiner
US 2022/0208355 A12022/0208355 A1 * 6/2022 Li......................... G06T 7/0016examiner
US 2022/0392428 A12022/0392428 A1 * 12/2022 Fernandez Guajardo...................examiner
A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms. K. Jeong, H.-K. Nguyen and H.-G. Kang, “A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms,” 2021 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021, pp. 41-45, doi: 10.23919/EUSIPCO54536.2021.9616247. (Year: 2021).10.23919/EUSIPCO54536.2021.9616247
IVCGAN:An Improved GAN for Voice Conversion. W. Zhao, W. Wang, J. Chai and J. Huang, “IVCGAN:An Improved GAN for Voice Conversion,” 2021 IEEE 5th Information Technology,Networking, Electronic and Automation Control Con- ference (ITNEC), Xi’an, China, 2021, pp. 1035-1039, doi: 10.1109/ITNEC52019.2021.9587053. (Year: 2021).10.1109/ITNEC52019.2021.9587053
A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms. K. Jeong, H.-K. Nguyen and H.-G. Kang, “A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms,” 2021 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021, pp. 41-45, doi: 10.23919/EUSIPCO54536.2021.9616247. (Year: 2021).10.23919/EUSIPCO54536.2021.9616247
IVCGAN:An Improved GAN for Voice Conversion. W. Zhao, W. Wang, J. Chai and J. Huang, “IVCGAN:An Improved GAN for Voice Conversion,” 2021 IEEE 5th Information Technology,Networking, Electronic and Automation Control Con- ference (ITNEC), Xi’an, China, 2021, pp. 1035-1039, doi: 10.1109/ITNEC52019.2021.9587053. (Year: 2021).10.1109/ITNEC52019.2021.9587053
A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms. K. Jeong, H.-K. Nguyen and H.-G. Kang, “A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms,” 2021 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021, pp. 41-45, doi: 10.23919/EUSIPCO54536.2021.9616247. (Year: 2021).10.23919/EUSIPCO54536.2021.9616247
IVCGAN:An Improved GAN for Voice Conversion. W. Zhao, W. Wang, J. Chai and J. Huang, “IVCGAN:An Improved GAN for Voice Conversion,” 2021 IEEE 5th Information Technology,Networking, Electronic and Automation Control Con- ference (ITNEC), Xi’an, China, 2021, pp. 1035-1039, doi: 10.1109/ITNEC52019.2021.9587053. (Year: 2021).10.1109/ITNEC52019.2021.9587053
A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms. K. Jeong, H.-K. Nguyen and H.-G. Kang, “A Fast and Lightweight Text-to-Speech Model with Spectrum and Waveform Alignment Algorithms,” 2021 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021, pp. 41-45, doi: 10.23919/EUSIPCO54536.2021.9616247. (Year: 2021).10.23919/EUSIPCO54536.2021.9616247
IVCGAN:An Improved GAN for Voice Conversion. W. Zhao, W. Wang, J. Chai and J. Huang, “IVCGAN:An Improved GAN for Voice Conversion,” 2021 IEEE 5th Information Technology,Networking, Electronic and Automation Control Con- ference (ITNEC), Xi’an, China, 2021, pp. 1035-1039, doi: 10.1109/ITNEC52019.2021.9587053. (Year: 2021).10.1109/ITNEC52019.2021.9587053