Pre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images

dc.authorid0000-0002-5213-8517
dc.authorid0000-0002-5364-6265
dc.authorid0000-0003-0569-098X
dc.contributor.authorSerin, Zafer
dc.contributor.authorYüzgeç, Uğur
dc.contributor.authorKarakuzu, Cihan
dc.date.accessioned2026-07-20T05:46:26Z
dc.date.issued2024
dc.departmentMeslek Yüksekokulları, Pazaryeri Meslek Yüksekokulu, Bilgisayar Teknolojileri Bölümü
dc.departmentFakülteler, Mühendislik Fakültesi, Bilgisayar Mühendisliği Bölümü
dc.departmentEnstitüler, Lisansüstü Eğitim Enstitüsü, Elektronik ve Bilgisayar Mühendisliği Ana Bilim Dalı
dc.description.abstractBu çalışmada, 2B görüntülerden 3B nesne üretimi alanında üretken çekişmeli ağların (GAN'lar) ve varyasyonel oto-kodlayıcıların (VAE'ler) özgün bir kombinasyonu olan 3D-VAE-GAN modellerine odaklanılmaktadır. Spesifik olarak, yapının VAE bileşenindeki potansiyel kodlayıcı ağlar olarak önceden eğitilmiş çeşitli evrişimli sinir ağlarının (CNN'ler) kullanımı araştırılmaktadır. Bu önceden eğitilmiş ağ modelleri DenseNet121, EfficientNetB0, RegNet16 ve ResNet18'dir. Ayrıca, tam bağlantılı (fully connected) bir katmana sahip standart bir CNN modeli ile tam bağlantılı bir katmanı olmayan bir CNN modeli de kullanılmıştır. Eğitim sürecini kolaylaştırmak için, GAN modelinin üretici ve ayırt edici ağlarında ikili çapraz entropi (binary cross-entropy) kayıp fonksiyonu kullanılırken, VAE modelinin kodlayıcı ağı için Kullback-Leibler ıraksaması (divergence) kullanılmıştır. Eğitim ve test aşamaları için, ShapeNet veri setinde yer alan sandalye kategorisine odaklanılmıştır. 3B ögelerin temsili hususunda, sinir ağı ve derin öğrenme prosedürleriyle son derece uyumlu olduğu kanıtlanan voksellere dayalı bir seçim kullanılmıştır. Yaklaşımımız, 3B nesneler üretmek için girdi olarak yalnızca tek bir 2B görüntü kullanmaktadır. Testler ve değerlendirmeler, VAE'nin kodlayıcı ağı kısmında önceden eğitilmiş ağların kullanımının oldukça başarılı sonuçlar verdiğini göstermiştir. Elde edilen ortalama Kullback-Leibler ıraksama değerleri sırasıyla RegNet16 için 1129.660, ResNet18 için 1219.067, EfficientNetB0 için 1352.815, tam bağlantılı katmanı olmayan CNN için 1538.489, DenseNet121 için 2893.807 ve tam bağlantılı katmana sahip CNN için 1696.749 olarak bulunmuştur. Önceden eğitilmiş RegNet16, diğer yöntemlerden daha üstün bir performans sergilemektedir.
dc.description.abstractIn this study, we focus on the 3D-VAE-GAN models, a novel combination of generative adversarial networks (GANs) and variational autoencoders (VAEs) in the field of 3D object generation from 2D images. Specifically, we explore the use of several pre-trained convolutional neural networks (CNNs) as potential encoder networks in the VAE component of the structure. These pre-trained network models are DenseNet121, EfficientNetB0, RegNet16, and ResNet18. Additionally, a standard CNN model with a fully connected layer and a CNN model without a fully connected layer were also used. To facilitate the training process, the binary cross-entropy loss function is used for the generator and discriminator networks of the GAN model, while the Kullback–Leibler divergence is utilized for the encoder network of the VAE model. For the training and testing stages, attention is directed toward the chair category contained within the ShapeNet dataset. With regard to the depiction of 3D items, the selection utilized is based on voxels, which prove highly compatible with the neural network and deep learning procedures. Our approach uses only one 2D image as the input for producing 3D objects. Tests and evaluations have shown that the use of pre-trained networks in the encoder network portion of the VAE yields very successful results. The average Kullback–Leibler divergence values obtained were found to be 1129.660 for RegNet16, 1219.067 for ResNet18, 1352.815 for EfficientNetB0, 1538.489 for CNN without a fully connected layer, 2893.807 for DenseNet121, and 1696.749 for CNN with a fully connected layer, respectively. The pre-trained RegNet16 outperforms other methods.
dc.identifier.citationSerin, Z., Yüzgeç, U., Karakuzu, C. (2024). Pre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images. In: Seyman, M.N. (eds) 2nd International Congress of Electrical and Computer Engineering . ICECENG 2023. EAI/Springer Innovations in Communication and Computing. Springer, Cham. https://doi.org/10.1007/978-3-031-52760-9_7
dc.identifier.doi10.1007/978-3-031-52760-9_7
dc.identifier.scopusqualityQ2
dc.identifier.urihttps://doi.org/10.1007/978-3-031-52760-9_7
dc.identifier.urihttps://hdl.handle.net/11552/9724
dc.indekslendigikaynakScopus
dc.institutionauthorSerin, Zafer
dc.institutionauthorYüzgeç, Uğur
dc.institutionauthorKarakuzu, Cihan
dc.language.isoen
dc.publisherSpringer
dc.relation.ispartof2nd International Congress of Electrical and Computer Engineering
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı ve Öğrenci
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.titlePre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images
dc.typeConference Object

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
ICECENG-978-3-031-52760-9.pdf
Boyut:
12.61 MB
Biçim:
Adobe Portable Document Format

Lisans paketi

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
license.txt
Boyut:
1.17 KB
Biçim:
Item-specific license agreed upon to submission
Açıklama: