Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIGS. 1A, 1B and 1C illustrate an overview of training- 50 and/or architecture-related aspects disclosed herein, in accordance with respective embodiments.
FIG. 2
FIG. 2 is a table of images to illustrated disentanglement 60 of semantic attributes, namely, smile from eyeglasses, in facial images in accordance with an …
FIG. 3
FIG. 3 is a pseudocode listing operations, in accordance with an embodiment.
FIG. 4
FIG. 4 is a table of images to illustrate semantic attribute 65 manipulation results using a controllable neural network, in accordance with an embodiment. B₂
FIG. 5
FIG. 5 is a table of images to illustrate semantic attribute disentanglement results using a controllable neural network, in accordance with an embodiment.
FIG. 6
FIG. 6 is a table of images visualizing directions found to compare three controllable neural networks, namely two previously known controllable neural …
FIG. 7
FIG. 7 is a table of images showing a comparison of attribute disentanglement results obtained by a previously known controllable neural network and a …
FIG. 8
FIG. 8 is a block diagram of a graphical user interface (GUI) to produce a synthesized image from a source image using a controllable GAN, in accordance with …
FIG. 9
FIG. 9 is a block diagram of a computer system, in accordance with an embodiment.
FIG. 10
FIGS. 10 and 11 are flowcharts showing operations, in accordance with respective embodiments herein. The present concept is best described through certain …
FIG. 11
FIG. 11 is a flowchart of operations 1100 in accordance with an embodiment herein. Operations 1100 are performed by a computing device such as device 910, 912, …
FIG. 12
FIG. 13
FIG. 14
FIG. 15
FIG. 16
FIG. 16 is a graphical represen- tation of a portion of storage device 106 storing dimensions of respective gradients 132 and 130 for semantic attributes k and …
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
3 independent · 16 dependent
1
Independent
A method comprising: generating a synthesized image (g(z')) using a generator (g) having a latent code (z) which generator g manipu-lates a target semantic attribute (k) in the synthesized image, wherein the generating comprises: discovering a semantically meaningful data direction at z for the target semantic attribute k, the semantically meaningful data direction identified from an auxiliary network classifying respective semantic attributes at z, including target semantic attribute k, and the auxiliary network sharing a latent space Z with generator g, and where z∈Z; defining z' by optimizing latent code z responsive to the semantically meaningful direction for the target seman-tic attribute k; and disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion at z of the target semantic attribute k, those important dimensions of respective semantically mean-ingful data directions of each other semantic attribute (m, m≠k) entangled with the target semantic attribute k, wherein a particular dimension is important based on its absolute gradient magnitude; and outputting the synthesized image.
2
Dependent← claim 1
The method of claim 1, wherein the auxiliary network comprises a set of binary classifiers, one for each semantic attribute that the generator is capable to manipulate, the auxiliary network co-trained to share the latent space Z with generator g.
3
Dependent← claim 1
The method of claim 1, wherein the respective seman-tically meaningful data directions comprise a respective dimensional data vector obtained from each individual clas-sifier, each vector comprising a direction and rate of the fastest increase in the individual classifier.
4
Dependent← claim 1
The method of claim 1, wherein the removing com-prises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the target seman-tic attribute k; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimension in the semantically meaningful data direction of the target seman-tic attribute k to a zero value.
6
Dependent← claim 1
The method of claim 1, comprising repeating opera-tions of discovering, defining and optimizing in respect of latent code z' to further manipulate sematic attribute k.
7
Dependent← claim 1
The method of claim 1, wherein the generating of the synthesized image manipulates a plurality of target semantic attributes in the synthesized image, and the method com-prises discovering respective semantically meaningful data directions at z for each of the plurality of the target semantic attributes as identified from the auxiliary network classify-ing each of the plurality of the target semantic attributes at z, and optimizing z in response to each of the semantically meaningful data directions.
8
Dependent← claim 1
The method of claim 1 comprising training another network model using the synthesized image.
9
Dependent← claim 1
The method of claim 1, comprising receive an input identifying the target semantic attribute to be controlled relative to a source image.
11
Dependent← claim 1
The method of claim 1 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
12
Dependent← claim 1
The method of claim 1, wherein the target semantic attribute comprises one of: a facial feature comprising age, gender, smile, or other facial feature; a pose effect; a makeup effect; a hair effect; a nail effect; a cosmetic surgery or dental effect comprising one of a rhinoplasty, a lift, blepharoplasty, an implant, otoplasty, teeth whitening, teeth straightening or other cosmetic surgery or dental effect; or an appliance effect comprising one of an eye appliance, a mouth appliance, an ear appliance or other appliance effect.
13
Independent
A computer implemented method comprising execut-ing the steps comprising: providing a generator and an auxiliary network sharing a latent space, the generator configured to generate syn-thesized images exhibiting semantic attributes and the auxiliary network comprising a plurality of semantic attribute classifiers including a semantic attribute clas-sifier for each semantic attribute to be controlled for generating a synthesized image from a source image, each semantic attribute classifier configured to classify a presence of one of the semantic attributes in images and provide a semantically meaningful direction for controlling the one of the semantic attributes in the synthesized images of the generator; and generating the synthesized image from the source image using the generator by applying a respective semantic attribute control to control a respective semantic attri-bute in the synthesized image, the respective semantic attribute control responsive to the semantically mean-ingful direction provided by the classifier associated with the respective semantic attribute; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude.
14
Dependent← claim 13
The method of claim 13, wherein the semantically meaningful direction comprises a gradient direction, and the instructions cause the computing device to compute a respective semantic attribute control from a respective gra-dient direction of the classifier associated with the respective semantic attribute.
15
Dependent← claim 13
The method of claim 13, wherein the removing comprises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the respective semantic attribute; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimen-sion in the semantically meaningful data direction of the respective semantic attribute to a zero value.
16
Dependent← claim 13
The method of claim 13 comprising receiving an input identifying at least one of the respective semantic attributes to be controlled relative to the source image.
17
Dependent← claim 13
The method of claim 13 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
18
Independent
A method comprising the steps of: providing an augmented reality (AR) interface to provide an AR experience, the AR interface configured to generate a synthesized image from a received image using a generator by applying a respective semantic attribute control input to control a respective semantic attribute in the synthesized image, the respective semantic attribute control input configured to apply a semantically meaningful direction provided by a clas-sifier associated with the semantic attribute, the gen-erator and classifier sharing a latent space; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude; and receiving the received image and providing the synthe-sized image for the AR experience.
19
Dependent← claim 18
The method of claim 18 comprising processing the synthesized image using an effects pipeline to simulate an effect and providing the synthesized image with the simu-lated effect for presenting in the AR interface. ∗ ∗ ∗ ∗ ∗
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2015/0254289 A12015/0254289 A1 * 9/2015 Junkergard........... G06F 16/283examiner
US 2022/0207786 A12022/0207786 A1 * 6/2022 Ren........................ G06V 40/10examiner
US 2023/0015253 A12023/0015253 A1 * 1/2023 Nie........................ G06V 10/82examiner
US 2023/0153606 A12023/0153606 A1 * 5/2023 Min......................... G06N 3/08
Patent
Atlas literature
Patent
US 12,633,063 B2
METHODS AND APPARATUS FOR DETERMINING AND USING CONTROLLABLE DIRECTIONS OF GAN SPACE
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIGS. 1A, 1B and 1C illustrate an overview of training- 50 and/or architecture-related aspects disclosed herein, in accordance with respective embodiments.
FIG. 2
FIG. 2 is a table of images to illustrated disentanglement 60 of semantic attributes, namely, smile from eyeglasses, in facial images in accordance with an …
FIG. 3
FIG. 3 is a pseudocode listing operations, in accordance with an embodiment.
FIG. 4
FIG. 4 is a table of images to illustrate semantic attribute 65 manipulation results using a controllable neural network, in accordance with an embodiment. B₂
FIG. 5
FIG. 5 is a table of images to illustrate semantic attribute disentanglement results using a controllable neural network, in accordance with an embodiment.
FIG. 6
FIG. 6 is a table of images visualizing directions found to compare three controllable neural networks, namely two previously known controllable neural …
FIG. 7
FIG. 7 is a table of images showing a comparison of attribute disentanglement results obtained by a previously known controllable neural network and a …
FIG. 8
FIG. 8 is a block diagram of a graphical user interface (GUI) to produce a synthesized image from a source image using a controllable GAN, in accordance with …
FIG. 9
FIG. 9 is a block diagram of a computer system, in accordance with an embodiment.
FIG. 10
FIGS. 10 and 11 are flowcharts showing operations, in accordance with respective embodiments herein. The present concept is best described through certain …
FIG. 11
FIG. 11 is a flowchart of operations 1100 in accordance with an embodiment herein. Operations 1100 are performed by a computing device such as device 910, 912, …
FIG. 12
FIG. 13
FIG. 14
FIG. 15
FIG. 16
FIG. 16 is a graphical represen- tation of a portion of storage device 106 storing dimensions of respective gradients 132 and 130 for semantic attributes k and …
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
3 independent · 16 dependent
1
Independent
A method comprising: generating a synthesized image (g(z')) using a generator (g) having a latent code (z) which generator g manipu-lates a target semantic attribute (k) in the synthesized image, wherein the generating comprises: discovering a semantically meaningful data direction at z for the target semantic attribute k, the semantically meaningful data direction identified from an auxiliary network classifying respective semantic attributes at z, including target semantic attribute k, and the auxiliary network sharing a latent space Z with generator g, and where z∈Z; defining z' by optimizing latent code z responsive to the semantically meaningful direction for the target seman-tic attribute k; and disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion at z of the target semantic attribute k, those important dimensions of respective semantically mean-ingful data directions of each other semantic attribute (m, m≠k) entangled with the target semantic attribute k, wherein a particular dimension is important based on its absolute gradient magnitude; and outputting the synthesized image.
2
Dependent← claim 1
The method of claim 1, wherein the auxiliary network comprises a set of binary classifiers, one for each semantic attribute that the generator is capable to manipulate, the auxiliary network co-trained to share the latent space Z with generator g.
3
Dependent← claim 1
The method of claim 1, wherein the respective seman-tically meaningful data directions comprise a respective dimensional data vector obtained from each individual clas-sifier, each vector comprising a direction and rate of the fastest increase in the individual classifier.
4
Dependent← claim 1
The method of claim 1, wherein the removing com-prises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the target seman-tic attribute k; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimension in the semantically meaningful data direction of the target seman-tic attribute k to a zero value.
6
Dependent← claim 1
The method of claim 1, comprising repeating opera-tions of discovering, defining and optimizing in respect of latent code z' to further manipulate sematic attribute k.
7
Dependent← claim 1
The method of claim 1, wherein the generating of the synthesized image manipulates a plurality of target semantic attributes in the synthesized image, and the method com-prises discovering respective semantically meaningful data directions at z for each of the plurality of the target semantic attributes as identified from the auxiliary network classify-ing each of the plurality of the target semantic attributes at z, and optimizing z in response to each of the semantically meaningful data directions.
8
Dependent← claim 1
The method of claim 1 comprising training another network model using the synthesized image.
9
Dependent← claim 1
The method of claim 1, comprising receive an input identifying the target semantic attribute to be controlled relative to a source image.
11
Dependent← claim 1
The method of claim 1 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
12
Dependent← claim 1
The method of claim 1, wherein the target semantic attribute comprises one of: a facial feature comprising age, gender, smile, or other facial feature; a pose effect; a makeup effect; a hair effect; a nail effect; a cosmetic surgery or dental effect comprising one of a rhinoplasty, a lift, blepharoplasty, an implant, otoplasty, teeth whitening, teeth straightening or other cosmetic surgery or dental effect; or an appliance effect comprising one of an eye appliance, a mouth appliance, an ear appliance or other appliance effect.
13
Independent
A computer implemented method comprising execut-ing the steps comprising: providing a generator and an auxiliary network sharing a latent space, the generator configured to generate syn-thesized images exhibiting semantic attributes and the auxiliary network comprising a plurality of semantic attribute classifiers including a semantic attribute clas-sifier for each semantic attribute to be controlled for generating a synthesized image from a source image, each semantic attribute classifier configured to classify a presence of one of the semantic attributes in images and provide a semantically meaningful direction for controlling the one of the semantic attributes in the synthesized images of the generator; and generating the synthesized image from the source image using the generator by applying a respective semantic attribute control to control a respective semantic attri-bute in the synthesized image, the respective semantic attribute control responsive to the semantically mean-ingful direction provided by the classifier associated with the respective semantic attribute; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude.
14
Dependent← claim 13
The method of claim 13, wherein the semantically meaningful direction comprises a gradient direction, and the instructions cause the computing device to compute a respective semantic attribute control from a respective gra-dient direction of the classifier associated with the respective semantic attribute.
15
Dependent← claim 13
The method of claim 13, wherein the removing comprises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the respective semantic attribute; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimen-sion in the semantically meaningful data direction of the respective semantic attribute to a zero value.
16
Dependent← claim 13
The method of claim 13 comprising receiving an input identifying at least one of the respective semantic attributes to be controlled relative to the source image.
17
Dependent← claim 13
The method of claim 13 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
18
Independent
A method comprising the steps of: providing an augmented reality (AR) interface to provide an AR experience, the AR interface configured to generate a synthesized image from a received image using a generator by applying a respective semantic attribute control input to control a respective semantic attribute in the synthesized image, the respective semantic attribute control input configured to apply a semantically meaningful direction provided by a clas-sifier associated with the semantic attribute, the gen-erator and classifier sharing a latent space; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude; and receiving the received image and providing the synthe-sized image for the AR experience.
19
Dependent← claim 18
The method of claim 18 comprising processing the synthesized image using an effects pipeline to simulate an effect and providing the synthesized image with the simu-lated effect for presenting in the AR interface. ∗ ∗ ∗ ∗ ∗
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2015/0254289 A12015/0254289 A1 * 9/2015 Junkergard........... G06F 16/283examiner
US 2022/0207786 A12022/0207786 A1 * 6/2022 Ren........................ G06V 40/10examiner
US 2023/0015253 A12023/0015253 A1 * 1/2023 Nie........................ G06V 10/82examiner
US 2023/0153606 A12023/0153606 A1 * 5/2023 Min......................... G06N 3/08
Patent
Atlas literature
Patent
US 12,633,063 B2
METHODS AND APPARATUS FOR DETERMINING AND USING CONTROLLABLE DIRECTIONS OF GAN SPACE
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIGS. 1A, 1B and 1C illustrate an overview of training- 50 and/or architecture-related aspects disclosed herein, in accordance with respective embodiments.
FIG. 2
FIG. 2 is a table of images to illustrated disentanglement 60 of semantic attributes, namely, smile from eyeglasses, in facial images in accordance with an …
FIG. 3
FIG. 3 is a pseudocode listing operations, in accordance with an embodiment.
FIG. 4
FIG. 4 is a table of images to illustrate semantic attribute 65 manipulation results using a controllable neural network, in accordance with an embodiment. B₂
FIG. 5
FIG. 5 is a table of images to illustrate semantic attribute disentanglement results using a controllable neural network, in accordance with an embodiment.
FIG. 6
FIG. 6 is a table of images visualizing directions found to compare three controllable neural networks, namely two previously known controllable neural …
FIG. 7
FIG. 7 is a table of images showing a comparison of attribute disentanglement results obtained by a previously known controllable neural network and a …
FIG. 8
FIG. 8 is a block diagram of a graphical user interface (GUI) to produce a synthesized image from a source image using a controllable GAN, in accordance with …
FIG. 9
FIG. 9 is a block diagram of a computer system, in accordance with an embodiment.
FIG. 10
FIGS. 10 and 11 are flowcharts showing operations, in accordance with respective embodiments herein. The present concept is best described through certain …
FIG. 11
FIG. 11 is a flowchart of operations 1100 in accordance with an embodiment herein. Operations 1100 are performed by a computing device such as device 910, 912, …
FIG. 12
FIG. 13
FIG. 14
FIG. 15
FIG. 16
FIG. 16 is a graphical represen- tation of a portion of storage device 106 storing dimensions of respective gradients 132 and 130 for semantic attributes k and …
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
3 independent · 16 dependent
1
Independent
A method comprising: generating a synthesized image (g(z')) using a generator (g) having a latent code (z) which generator g manipu-lates a target semantic attribute (k) in the synthesized image, wherein the generating comprises: discovering a semantically meaningful data direction at z for the target semantic attribute k, the semantically meaningful data direction identified from an auxiliary network classifying respective semantic attributes at z, including target semantic attribute k, and the auxiliary network sharing a latent space Z with generator g, and where z∈Z; defining z' by optimizing latent code z responsive to the semantically meaningful direction for the target seman-tic attribute k; and disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion at z of the target semantic attribute k, those important dimensions of respective semantically mean-ingful data directions of each other semantic attribute (m, m≠k) entangled with the target semantic attribute k, wherein a particular dimension is important based on its absolute gradient magnitude; and outputting the synthesized image.
2
Dependent← claim 1
The method of claim 1, wherein the auxiliary network comprises a set of binary classifiers, one for each semantic attribute that the generator is capable to manipulate, the auxiliary network co-trained to share the latent space Z with generator g.
3
Dependent← claim 1
The method of claim 1, wherein the respective seman-tically meaningful data directions comprise a respective dimensional data vector obtained from each individual clas-sifier, each vector comprising a direction and rate of the fastest increase in the individual classifier.
4
Dependent← claim 1
The method of claim 1, wherein the removing com-prises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the target seman-tic attribute k; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimension in the semantically meaningful data direction of the target seman-tic attribute k to a zero value.
6
Dependent← claim 1
The method of claim 1, comprising repeating opera-tions of discovering, defining and optimizing in respect of latent code z' to further manipulate sematic attribute k.
7
Dependent← claim 1
The method of claim 1, wherein the generating of the synthesized image manipulates a plurality of target semantic attributes in the synthesized image, and the method com-prises discovering respective semantically meaningful data directions at z for each of the plurality of the target semantic attributes as identified from the auxiliary network classify-ing each of the plurality of the target semantic attributes at z, and optimizing z in response to each of the semantically meaningful data directions.
8
Dependent← claim 1
The method of claim 1 comprising training another network model using the synthesized image.
9
Dependent← claim 1
The method of claim 1, comprising receive an input identifying the target semantic attribute to be controlled relative to a source image.
11
Dependent← claim 1
The method of claim 1 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
12
Dependent← claim 1
The method of claim 1, wherein the target semantic attribute comprises one of: a facial feature comprising age, gender, smile, or other facial feature; a pose effect; a makeup effect; a hair effect; a nail effect; a cosmetic surgery or dental effect comprising one of a rhinoplasty, a lift, blepharoplasty, an implant, otoplasty, teeth whitening, teeth straightening or other cosmetic surgery or dental effect; or an appliance effect comprising one of an eye appliance, a mouth appliance, an ear appliance or other appliance effect.
13
Independent
A computer implemented method comprising execut-ing the steps comprising: providing a generator and an auxiliary network sharing a latent space, the generator configured to generate syn-thesized images exhibiting semantic attributes and the auxiliary network comprising a plurality of semantic attribute classifiers including a semantic attribute clas-sifier for each semantic attribute to be controlled for generating a synthesized image from a source image, each semantic attribute classifier configured to classify a presence of one of the semantic attributes in images and provide a semantically meaningful direction for controlling the one of the semantic attributes in the synthesized images of the generator; and generating the synthesized image from the source image using the generator by applying a respective semantic attribute control to control a respective semantic attri-bute in the synthesized image, the respective semantic attribute control responsive to the semantically mean-ingful direction provided by the classifier associated with the respective semantic attribute; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude.
14
Dependent← claim 13
The method of claim 13, wherein the semantically meaningful direction comprises a gradient direction, and the instructions cause the computing device to compute a respective semantic attribute control from a respective gra-dient direction of the classifier associated with the respective semantic attribute.
15
Dependent← claim 13
The method of claim 13, wherein the removing comprises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the respective semantic attribute; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimen-sion in the semantically meaningful data direction of the respective semantic attribute to a zero value.
16
Dependent← claim 13
The method of claim 13 comprising receiving an input identifying at least one of the respective semantic attributes to be controlled relative to the source image.
17
Dependent← claim 13
The method of claim 13 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
18
Independent
A method comprising the steps of: providing an augmented reality (AR) interface to provide an AR experience, the AR interface configured to generate a synthesized image from a received image using a generator by applying a respective semantic attribute control input to control a respective semantic attribute in the synthesized image, the respective semantic attribute control input configured to apply a semantically meaningful direction provided by a clas-sifier associated with the semantic attribute, the gen-erator and classifier sharing a latent space; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude; and receiving the received image and providing the synthe-sized image for the AR experience.
19
Dependent← claim 18
The method of claim 18 comprising processing the synthesized image using an effects pipeline to simulate an effect and providing the synthesized image with the simu-lated effect for presenting in the AR interface. ∗ ∗ ∗ ∗ ∗
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2015/0254289 A12015/0254289 A1 * 9/2015 Junkergard........... G06F 16/283examiner
US 2022/0207786 A12022/0207786 A1 * 6/2022 Ren........................ G06V 40/10examiner
US 2023/0015253 A12023/0015253 A1 * 1/2023 Nie........................ G06V 10/82examiner
US 2023/0153606 A12023/0153606 A1 * 5/2023 Min......................... G06N 3/08
Patent
Atlas literature
Patent
US 12,633,063 B2
METHODS AND APPARATUS FOR DETERMINING AND USING CONTROLLABLE DIRECTIONS OF GAN SPACE
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIGS. 1A, 1B and 1C illustrate an overview of training- 50 and/or architecture-related aspects disclosed herein, in accordance with respective embodiments.
FIG. 2
FIG. 2 is a table of images to illustrated disentanglement 60 of semantic attributes, namely, smile from eyeglasses, in facial images in accordance with an …
FIG. 3
FIG. 3 is a pseudocode listing operations, in accordance with an embodiment.
FIG. 4
FIG. 4 is a table of images to illustrate semantic attribute 65 manipulation results using a controllable neural network, in accordance with an embodiment. B₂
FIG. 5
FIG. 5 is a table of images to illustrate semantic attribute disentanglement results using a controllable neural network, in accordance with an embodiment.
FIG. 6
FIG. 6 is a table of images visualizing directions found to compare three controllable neural networks, namely two previously known controllable neural …
FIG. 7
FIG. 7 is a table of images showing a comparison of attribute disentanglement results obtained by a previously known controllable neural network and a …
FIG. 8
FIG. 8 is a block diagram of a graphical user interface (GUI) to produce a synthesized image from a source image using a controllable GAN, in accordance with …
FIG. 9
FIG. 9 is a block diagram of a computer system, in accordance with an embodiment.
FIG. 10
FIGS. 10 and 11 are flowcharts showing operations, in accordance with respective embodiments herein. The present concept is best described through certain …
FIG. 11
FIG. 11 is a flowchart of operations 1100 in accordance with an embodiment herein. Operations 1100 are performed by a computing device such as device 910, 912, …
FIG. 12
FIG. 13
FIG. 14
FIG. 15
FIG. 16
FIG. 16 is a graphical represen- tation of a portion of storage device 106 storing dimensions of respective gradients 132 and 130 for semantic attributes k and …
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
3 independent · 16 dependent
1
Independent
A method comprising: generating a synthesized image (g(z')) using a generator (g) having a latent code (z) which generator g manipu-lates a target semantic attribute (k) in the synthesized image, wherein the generating comprises: discovering a semantically meaningful data direction at z for the target semantic attribute k, the semantically meaningful data direction identified from an auxiliary network classifying respective semantic attributes at z, including target semantic attribute k, and the auxiliary network sharing a latent space Z with generator g, and where z∈Z; defining z' by optimizing latent code z responsive to the semantically meaningful direction for the target seman-tic attribute k; and disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion at z of the target semantic attribute k, those important dimensions of respective semantically mean-ingful data directions of each other semantic attribute (m, m≠k) entangled with the target semantic attribute k, wherein a particular dimension is important based on its absolute gradient magnitude; and outputting the synthesized image.
2
Dependent← claim 1
The method of claim 1, wherein the auxiliary network comprises a set of binary classifiers, one for each semantic attribute that the generator is capable to manipulate, the auxiliary network co-trained to share the latent space Z with generator g.
3
Dependent← claim 1
The method of claim 1, wherein the respective seman-tically meaningful data directions comprise a respective dimensional data vector obtained from each individual clas-sifier, each vector comprising a direction and rate of the fastest increase in the individual classifier.
4
Dependent← claim 1
The method of claim 1, wherein the removing com-prises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the target seman-tic attribute k; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimension in the semantically meaningful data direction of the target seman-tic attribute k to a zero value.
6
Dependent← claim 1
The method of claim 1, comprising repeating opera-tions of discovering, defining and optimizing in respect of latent code z' to further manipulate sematic attribute k.
7
Dependent← claim 1
The method of claim 1, wherein the generating of the synthesized image manipulates a plurality of target semantic attributes in the synthesized image, and the method com-prises discovering respective semantically meaningful data directions at z for each of the plurality of the target semantic attributes as identified from the auxiliary network classify-ing each of the plurality of the target semantic attributes at z, and optimizing z in response to each of the semantically meaningful data directions.
8
Dependent← claim 1
The method of claim 1 comprising training another network model using the synthesized image.
9
Dependent← claim 1
The method of claim 1, comprising receive an input identifying the target semantic attribute to be controlled relative to a source image.
11
Dependent← claim 1
The method of claim 1 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
12
Dependent← claim 1
The method of claim 1, wherein the target semantic attribute comprises one of: a facial feature comprising age, gender, smile, or other facial feature; a pose effect; a makeup effect; a hair effect; a nail effect; a cosmetic surgery or dental effect comprising one of a rhinoplasty, a lift, blepharoplasty, an implant, otoplasty, teeth whitening, teeth straightening or other cosmetic surgery or dental effect; or an appliance effect comprising one of an eye appliance, a mouth appliance, an ear appliance or other appliance effect.
13
Independent
A computer implemented method comprising execut-ing the steps comprising: providing a generator and an auxiliary network sharing a latent space, the generator configured to generate syn-thesized images exhibiting semantic attributes and the auxiliary network comprising a plurality of semantic attribute classifiers including a semantic attribute clas-sifier for each semantic attribute to be controlled for generating a synthesized image from a source image, each semantic attribute classifier configured to classify a presence of one of the semantic attributes in images and provide a semantically meaningful direction for controlling the one of the semantic attributes in the synthesized images of the generator; and generating the synthesized image from the source image using the generator by applying a respective semantic attribute control to control a respective semantic attri-bute in the synthesized image, the respective semantic attribute control responsive to the semantically mean-ingful direction provided by the classifier associated with the respective semantic attribute; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude.
14
Dependent← claim 13
The method of claim 13, wherein the semantically meaningful direction comprises a gradient direction, and the instructions cause the computing device to compute a respective semantic attribute control from a respective gra-dient direction of the classifier associated with the respective semantic attribute.
15
Dependent← claim 13
The method of claim 13, wherein the removing comprises evaluating each dimension of each semantically meaningful data direction to be disentangled from the semantically meaningful data direction of the respective semantic attribute; if a particular dimension exceeds a threshold value, setting a value of the corresponding dimen-sion in the semantically meaningful data direction of the respective semantic attribute to a zero value.
16
Dependent← claim 13
The method of claim 13 comprising receiving an input identifying at least one of the respective semantic attributes to be controlled relative to the source image.
17
Dependent← claim 13
The method of claim 13 comprising any one or more of: providing the generator and the auxiliary network to generate the synthesized images as a service; providing an e-commerce interface to purchase a product or service; providing a recommendation interface to recommend a product or service; or providing an augmented reality interface using the syn-thesized image to provide an augmented reality expe-rience.
18
Independent
A method comprising the steps of: providing an augmented reality (AR) interface to provide an AR experience, the AR interface configured to generate a synthesized image from a received image using a generator by applying a respective semantic attribute control input to control a respective semantic attribute in the synthesized image, the respective semantic attribute control input configured to apply a semantically meaningful direction provided by a clas-sifier associated with the semantic attribute, the gen-erator and classifier sharing a latent space; wherein to generate the synthesized image uses a latent code z' defined by optimizing a latent code z of the latent space responsive to the semantically meaningful direction for the respective semantic attribute, includ-ing disentangling data directions by removing out, from dimensions of the semantically meaningful data direc-tion of the respective semantic attribute, those impor-tant dimensions of respective semantically meaningful data directions of each other semantic attribute entangled with the respective semantic attribute, wherein a particular dimension is important based on its absolute gradient magnitude; and receiving the received image and providing the synthe-sized image for the AR experience.
19
Dependent← claim 18
The method of claim 18 comprising processing the synthesized image using an effects pipeline to simulate an effect and providing the synthesized image with the simu-lated effect for presenting in the AR interface. ∗ ∗ ∗ ∗ ∗
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2015/0254289 A12015/0254289 A1 * 9/2015 Junkergard........... G06F 16/283examiner
US 2022/0207786 A12022/0207786 A1 * 6/2022 Ren........................ G06V 40/10examiner
US 2023/0015253 A12023/0015253 A1 * 1/2023 Nie........................ G06V 10/82examiner
US 2023/0153606 A12023/0153606 A1 * 5/2023 Min......................... G06N 3/08
7
examiner
US 2023/0252692 A12023/0252692 A1 * 8/2023 Liu........................... G06T 3/18examiner
Cited non-patent literature · 2
Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes. Huiting Yang et al. “Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes”, 2021, IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp. 12177- 12185(OpenAccessVersion,providedbyComputerVisionFoundation) (Year: 2021).
Wasserstein generative adversarial net- works. Martin Arjovsky et al., “Wasserstein generative adversarial net- works”, In International conference on machine learning, pp. 214- 223, PMLR, 2017. Andrew Brock et al., “Large scale GAN training for high fidelity natural image synthesis”, arXiv preprint arXiv:1809, 11096, pp. 1-35, 2018. Yunjey Choi et al., “StarGAN: Unified generative adversarial net- works for multi-domain image-to-image translation”, In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pp. 8789-8797, 2018. Anton Cherepkov et al., “Navigating the GAN parameter space for semantic image editing”, In Proceedings of the IEEEICVF Confer- ence on Computer Vision and Pattern Recognition, pp. 3671-3680, 2021. Ian Goodfellow et al., “Generative adversarial nets”, Advances in neural information processing systems, pp. 1-9, 27, 2014. Erik Ha¨rko¨nen et al., “GANSpace: Discovering interpretable GAN controls”, atX1v preprint atXiv:2004. 02546, 2020, pp. 1-10. Yujun Shen et al., “InterFaceGan: Interpreting the disentangled face representation learned by GANs”, IEEE transactions on pattern analysis and machine intelligence, XP93093638, vol. 44, No. 4, 2020, pp. 2004-2018. Yujun Shen et al., “Closed-form factorization of latent semantics in GANs”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1532-1540, 2021. Andrey Voynov et al., “Unsupervised discovery of interpretable directions in the GAN latent space”, In International Conference on Machine Learning, pp. 9786-9796, PMLR, 2020. Yan Wu et al., “Logan: Latent optimisation for generative adversarial networks”, arXiv preprint arXiv:1912.00953, 2019, pp. 1-23. Zongze Wu et al., “StyleSpace analysis; Disentangled controls for styleGAN image generation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12863-12872, 2021. Bolei Zhou et al., “Learning deep features for discriminative localization”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2921-2929, 2016. International Search Report and Written Opinion issued Nov. 2, 2023, in PCT/EP2023/070911, 20 pages. Robin Kips et al., “CA-GAN: Weakly Supervised Color Aware GAN for Controllable Makeup Transfer”, Arxiv.Org, Cornell Uni- versity Library, 201 Olin Library Cornell University Ithaca, NY 14853, XP81890729, 16 pages. Https://www.Youtube.com/@digitalsreeni: “257—Exploring GAN latent space to generate images with desired features?”, Feb. 16, 2022, XP093094048, 2 pages Retrieved from the Internet: URL:https://www.youtube.com/watch?v=iuQ_f3W5Ttk [retrieved on Oct. 23, 2023]. Written Opinion and Search Report dated May 3, 2023 in French Application No. 2210849. Zikun Chen et al. “Exploring Gradient-based Multi-directional Controls in GANS.” Sep. 1, 2022. pp. 1-17. XP093043623; URL:https://arxiv.org/pdf/2209.00698.pdf. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi-librium. Advances in neural information processing systems, 30, 2017. Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Alla. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119, 2020. Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:1812.04948, 2018. Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pp. 3730-3738, 2015. Antoine Plumerault, Herve´ Le Borgne, and Ce´line Hudelot. Con- trolling generative models with continuous factors of variations. arXiv preprint arXiv:2001.10238, 2020. Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based local- ization. In Proceedings of the IEEE international conference on computer vision, pp. 618-626, 2017. Japanese Office Action dated Sep. 17, 2025, issued in Japanese Patent Application No. 2025-504163 (with English translation).
7
examiner
US 2023/0252692 A12023/0252692 A1 * 8/2023 Liu........................... G06T 3/18examiner
Cited non-patent literature · 2
Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes. Huiting Yang et al. “Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes”, 2021, IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp. 12177- 12185(OpenAccessVersion,providedbyComputerVisionFoundation) (Year: 2021).
Wasserstein generative adversarial net- works. Martin Arjovsky et al., “Wasserstein generative adversarial net- works”, In International conference on machine learning, pp. 214- 223, PMLR, 2017. Andrew Brock et al., “Large scale GAN training for high fidelity natural image synthesis”, arXiv preprint arXiv:1809, 11096, pp. 1-35, 2018. Yunjey Choi et al., “StarGAN: Unified generative adversarial net- works for multi-domain image-to-image translation”, In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pp. 8789-8797, 2018. Anton Cherepkov et al., “Navigating the GAN parameter space for semantic image editing”, In Proceedings of the IEEEICVF Confer- ence on Computer Vision and Pattern Recognition, pp. 3671-3680, 2021. Ian Goodfellow et al., “Generative adversarial nets”, Advances in neural information processing systems, pp. 1-9, 27, 2014. Erik Ha¨rko¨nen et al., “GANSpace: Discovering interpretable GAN controls”, atX1v preprint atXiv:2004. 02546, 2020, pp. 1-10. Yujun Shen et al., “InterFaceGan: Interpreting the disentangled face representation learned by GANs”, IEEE transactions on pattern analysis and machine intelligence, XP93093638, vol. 44, No. 4, 2020, pp. 2004-2018. Yujun Shen et al., “Closed-form factorization of latent semantics in GANs”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1532-1540, 2021. Andrey Voynov et al., “Unsupervised discovery of interpretable directions in the GAN latent space”, In International Conference on Machine Learning, pp. 9786-9796, PMLR, 2020. Yan Wu et al., “Logan: Latent optimisation for generative adversarial networks”, arXiv preprint arXiv:1912.00953, 2019, pp. 1-23. Zongze Wu et al., “StyleSpace analysis; Disentangled controls for styleGAN image generation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12863-12872, 2021. Bolei Zhou et al., “Learning deep features for discriminative localization”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2921-2929, 2016. International Search Report and Written Opinion issued Nov. 2, 2023, in PCT/EP2023/070911, 20 pages. Robin Kips et al., “CA-GAN: Weakly Supervised Color Aware GAN for Controllable Makeup Transfer”, Arxiv.Org, Cornell Uni- versity Library, 201 Olin Library Cornell University Ithaca, NY 14853, XP81890729, 16 pages. Https://www.Youtube.com/@digitalsreeni: “257—Exploring GAN latent space to generate images with desired features?”, Feb. 16, 2022, XP093094048, 2 pages Retrieved from the Internet: URL:https://www.youtube.com/watch?v=iuQ_f3W5Ttk [retrieved on Oct. 23, 2023]. Written Opinion and Search Report dated May 3, 2023 in French Application No. 2210849. Zikun Chen et al. “Exploring Gradient-based Multi-directional Controls in GANS.” Sep. 1, 2022. pp. 1-17. XP093043623; URL:https://arxiv.org/pdf/2209.00698.pdf. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi-librium. Advances in neural information processing systems, 30, 2017. Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Alla. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119, 2020. Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:1812.04948, 2018. Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pp. 3730-3738, 2015. Antoine Plumerault, Herve´ Le Borgne, and Ce´line Hudelot. Con- trolling generative models with continuous factors of variations. arXiv preprint arXiv:2001.10238, 2020. Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based local- ization. In Proceedings of the IEEE international conference on computer vision, pp. 618-626, 2017. Japanese Office Action dated Sep. 17, 2025, issued in Japanese Patent Application No. 2025-504163 (with English translation).
7
examiner
US 2023/0252692 A12023/0252692 A1 * 8/2023 Liu........................... G06T 3/18examiner
Cited non-patent literature · 2
Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes. Huiting Yang et al. “Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes”, 2021, IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp. 12177- 12185(OpenAccessVersion,providedbyComputerVisionFoundation) (Year: 2021).
Wasserstein generative adversarial net- works. Martin Arjovsky et al., “Wasserstein generative adversarial net- works”, In International conference on machine learning, pp. 214- 223, PMLR, 2017. Andrew Brock et al., “Large scale GAN training for high fidelity natural image synthesis”, arXiv preprint arXiv:1809, 11096, pp. 1-35, 2018. Yunjey Choi et al., “StarGAN: Unified generative adversarial net- works for multi-domain image-to-image translation”, In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pp. 8789-8797, 2018. Anton Cherepkov et al., “Navigating the GAN parameter space for semantic image editing”, In Proceedings of the IEEEICVF Confer- ence on Computer Vision and Pattern Recognition, pp. 3671-3680, 2021. Ian Goodfellow et al., “Generative adversarial nets”, Advances in neural information processing systems, pp. 1-9, 27, 2014. Erik Ha¨rko¨nen et al., “GANSpace: Discovering interpretable GAN controls”, atX1v preprint atXiv:2004. 02546, 2020, pp. 1-10. Yujun Shen et al., “InterFaceGan: Interpreting the disentangled face representation learned by GANs”, IEEE transactions on pattern analysis and machine intelligence, XP93093638, vol. 44, No. 4, 2020, pp. 2004-2018. Yujun Shen et al., “Closed-form factorization of latent semantics in GANs”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1532-1540, 2021. Andrey Voynov et al., “Unsupervised discovery of interpretable directions in the GAN latent space”, In International Conference on Machine Learning, pp. 9786-9796, PMLR, 2020. Yan Wu et al., “Logan: Latent optimisation for generative adversarial networks”, arXiv preprint arXiv:1912.00953, 2019, pp. 1-23. Zongze Wu et al., “StyleSpace analysis; Disentangled controls for styleGAN image generation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12863-12872, 2021. Bolei Zhou et al., “Learning deep features for discriminative localization”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2921-2929, 2016. International Search Report and Written Opinion issued Nov. 2, 2023, in PCT/EP2023/070911, 20 pages. Robin Kips et al., “CA-GAN: Weakly Supervised Color Aware GAN for Controllable Makeup Transfer”, Arxiv.Org, Cornell Uni- versity Library, 201 Olin Library Cornell University Ithaca, NY 14853, XP81890729, 16 pages. Https://www.Youtube.com/@digitalsreeni: “257—Exploring GAN latent space to generate images with desired features?”, Feb. 16, 2022, XP093094048, 2 pages Retrieved from the Internet: URL:https://www.youtube.com/watch?v=iuQ_f3W5Ttk [retrieved on Oct. 23, 2023]. Written Opinion and Search Report dated May 3, 2023 in French Application No. 2210849. Zikun Chen et al. “Exploring Gradient-based Multi-directional Controls in GANS.” Sep. 1, 2022. pp. 1-17. XP093043623; URL:https://arxiv.org/pdf/2209.00698.pdf. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi-librium. Advances in neural information processing systems, 30, 2017. Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Alla. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119, 2020. Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:1812.04948, 2018. Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pp. 3730-3738, 2015. Antoine Plumerault, Herve´ Le Borgne, and Ce´line Hudelot. Con- trolling generative models with continuous factors of variations. arXiv preprint arXiv:2001.10238, 2020. Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based local- ization. In Proceedings of the IEEE international conference on computer vision, pp. 618-626, 2017. Japanese Office Action dated Sep. 17, 2025, issued in Japanese Patent Application No. 2025-504163 (with English translation).
7
examiner
US 2023/0252692 A12023/0252692 A1 * 8/2023 Liu........................... G06T 3/18examiner
Cited non-patent literature · 2
Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes. Huiting Yang et al. “Discovering Interpretable Latent Space Direc- tions of GANs Beyond Binary Attributes”, 2021, IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp. 12177- 12185(OpenAccessVersion,providedbyComputerVisionFoundation) (Year: 2021).
Wasserstein generative adversarial net- works. Martin Arjovsky et al., “Wasserstein generative adversarial net- works”, In International conference on machine learning, pp. 214- 223, PMLR, 2017. Andrew Brock et al., “Large scale GAN training for high fidelity natural image synthesis”, arXiv preprint arXiv:1809, 11096, pp. 1-35, 2018. Yunjey Choi et al., “StarGAN: Unified generative adversarial net- works for multi-domain image-to-image translation”, In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pp. 8789-8797, 2018. Anton Cherepkov et al., “Navigating the GAN parameter space for semantic image editing”, In Proceedings of the IEEEICVF Confer- ence on Computer Vision and Pattern Recognition, pp. 3671-3680, 2021. Ian Goodfellow et al., “Generative adversarial nets”, Advances in neural information processing systems, pp. 1-9, 27, 2014. Erik Ha¨rko¨nen et al., “GANSpace: Discovering interpretable GAN controls”, atX1v preprint atXiv:2004. 02546, 2020, pp. 1-10. Yujun Shen et al., “InterFaceGan: Interpreting the disentangled face representation learned by GANs”, IEEE transactions on pattern analysis and machine intelligence, XP93093638, vol. 44, No. 4, 2020, pp. 2004-2018. Yujun Shen et al., “Closed-form factorization of latent semantics in GANs”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1532-1540, 2021. Andrey Voynov et al., “Unsupervised discovery of interpretable directions in the GAN latent space”, In International Conference on Machine Learning, pp. 9786-9796, PMLR, 2020. Yan Wu et al., “Logan: Latent optimisation for generative adversarial networks”, arXiv preprint arXiv:1912.00953, 2019, pp. 1-23. Zongze Wu et al., “StyleSpace analysis; Disentangled controls for styleGAN image generation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12863-12872, 2021. Bolei Zhou et al., “Learning deep features for discriminative localization”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2921-2929, 2016. International Search Report and Written Opinion issued Nov. 2, 2023, in PCT/EP2023/070911, 20 pages. Robin Kips et al., “CA-GAN: Weakly Supervised Color Aware GAN for Controllable Makeup Transfer”, Arxiv.Org, Cornell Uni- versity Library, 201 Olin Library Cornell University Ithaca, NY 14853, XP81890729, 16 pages. Https://www.Youtube.com/@digitalsreeni: “257—Exploring GAN latent space to generate images with desired features?”, Feb. 16, 2022, XP093094048, 2 pages Retrieved from the Internet: URL:https://www.youtube.com/watch?v=iuQ_f3W5Ttk [retrieved on Oct. 23, 2023]. Written Opinion and Search Report dated May 3, 2023 in French Application No. 2210849. Zikun Chen et al. “Exploring Gradient-based Multi-directional Controls in GANS.” Sep. 1, 2022. pp. 1-17. XP093043623; URL:https://arxiv.org/pdf/2209.00698.pdf. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi-librium. Advances in neural information processing systems, 30, 2017. Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Alla. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110-8119, 2020. Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:1812.04948, 2018. Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pp. 3730-3738, 2015. Antoine Plumerault, Herve´ Le Borgne, and Ce´line Hudelot. Con- trolling generative models with continuous factors of variations. arXiv preprint arXiv:2001.10238, 2020. Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based local- ization. In Proceedings of the IEEE international conference on computer vision, pp. 618-626, 2017. Japanese Office Action dated Sep. 17, 2025, issued in Japanese Patent Application No. 2025-504163 (with English translation).