Pre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images
| dc.authorid | 0000-0002-5213-8517 | |
| dc.authorid | 0000-0002-5364-6265 | |
| dc.authorid | 0000-0003-0569-098X | |
| dc.contributor.author | Serin, Zafer | |
| dc.contributor.author | Yüzgeç, Uğur | |
| dc.contributor.author | Karakuzu, Cihan | |
| dc.date.accessioned | 2026-07-20T05:46:26Z | |
| dc.date.issued | 2024 | |
| dc.department | Meslek Yüksekokulları, Pazaryeri Meslek Yüksekokulu, Bilgisayar Teknolojileri Bölümü | |
| dc.department | Fakülteler, Mühendislik Fakültesi, Bilgisayar Mühendisliği Bölümü | |
| dc.department | Enstitüler, Lisansüstü Eğitim Enstitüsü, Elektronik ve Bilgisayar Mühendisliği Ana Bilim Dalı | |
| dc.description.abstract | Bu çalışmada, 2B görüntülerden 3B nesne üretimi alanında üretken çekişmeli ağların (GAN'lar) ve varyasyonel oto-kodlayıcıların (VAE'ler) özgün bir kombinasyonu olan 3D-VAE-GAN modellerine odaklanılmaktadır. Spesifik olarak, yapının VAE bileşenindeki potansiyel kodlayıcı ağlar olarak önceden eğitilmiş çeşitli evrişimli sinir ağlarının (CNN'ler) kullanımı araştırılmaktadır. Bu önceden eğitilmiş ağ modelleri DenseNet121, EfficientNetB0, RegNet16 ve ResNet18'dir. Ayrıca, tam bağlantılı (fully connected) bir katmana sahip standart bir CNN modeli ile tam bağlantılı bir katmanı olmayan bir CNN modeli de kullanılmıştır. Eğitim sürecini kolaylaştırmak için, GAN modelinin üretici ve ayırt edici ağlarında ikili çapraz entropi (binary cross-entropy) kayıp fonksiyonu kullanılırken, VAE modelinin kodlayıcı ağı için Kullback-Leibler ıraksaması (divergence) kullanılmıştır. Eğitim ve test aşamaları için, ShapeNet veri setinde yer alan sandalye kategorisine odaklanılmıştır. 3B ögelerin temsili hususunda, sinir ağı ve derin öğrenme prosedürleriyle son derece uyumlu olduğu kanıtlanan voksellere dayalı bir seçim kullanılmıştır. Yaklaşımımız, 3B nesneler üretmek için girdi olarak yalnızca tek bir 2B görüntü kullanmaktadır. Testler ve değerlendirmeler, VAE'nin kodlayıcı ağı kısmında önceden eğitilmiş ağların kullanımının oldukça başarılı sonuçlar verdiğini göstermiştir. Elde edilen ortalama Kullback-Leibler ıraksama değerleri sırasıyla RegNet16 için 1129.660, ResNet18 için 1219.067, EfficientNetB0 için 1352.815, tam bağlantılı katmanı olmayan CNN için 1538.489, DenseNet121 için 2893.807 ve tam bağlantılı katmana sahip CNN için 1696.749 olarak bulunmuştur. Önceden eğitilmiş RegNet16, diğer yöntemlerden daha üstün bir performans sergilemektedir. | |
| dc.description.abstract | In this study, we focus on the 3D-VAE-GAN models, a novel combination of generative adversarial networks (GANs) and variational autoencoders (VAEs) in the field of 3D object generation from 2D images. Specifically, we explore the use of several pre-trained convolutional neural networks (CNNs) as potential encoder networks in the VAE component of the structure. These pre-trained network models are DenseNet121, EfficientNetB0, RegNet16, and ResNet18. Additionally, a standard CNN model with a fully connected layer and a CNN model without a fully connected layer were also used. To facilitate the training process, the binary cross-entropy loss function is used for the generator and discriminator networks of the GAN model, while the Kullback–Leibler divergence is utilized for the encoder network of the VAE model. For the training and testing stages, attention is directed toward the chair category contained within the ShapeNet dataset. With regard to the depiction of 3D items, the selection utilized is based on voxels, which prove highly compatible with the neural network and deep learning procedures. Our approach uses only one 2D image as the input for producing 3D objects. Tests and evaluations have shown that the use of pre-trained networks in the encoder network portion of the VAE yields very successful results. The average Kullback–Leibler divergence values obtained were found to be 1129.660 for RegNet16, 1219.067 for ResNet18, 1352.815 for EfficientNetB0, 1538.489 for CNN without a fully connected layer, 2893.807 for DenseNet121, and 1696.749 for CNN with a fully connected layer, respectively. The pre-trained RegNet16 outperforms other methods. | |
| dc.identifier.citation | Serin, Z., Yüzgeç, U., Karakuzu, C. (2024). Pre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images. In: Seyman, M.N. (eds) 2nd International Congress of Electrical and Computer Engineering . ICECENG 2023. EAI/Springer Innovations in Communication and Computing. Springer, Cham. https://doi.org/10.1007/978-3-031-52760-9_7 | |
| dc.identifier.doi | 10.1007/978-3-031-52760-9_7 | |
| dc.identifier.scopusquality | Q2 | |
| dc.identifier.uri | https://doi.org/10.1007/978-3-031-52760-9_7 | |
| dc.identifier.uri | https://hdl.handle.net/11552/9724 | |
| dc.indekslendigikaynak | Scopus | |
| dc.institutionauthor | Serin, Zafer | |
| dc.institutionauthor | Yüzgeç, Uğur | |
| dc.institutionauthor | Karakuzu, Cihan | |
| dc.language.iso | en | |
| dc.publisher | Springer | |
| dc.relation.ispartof | 2nd International Congress of Electrical and Computer Engineering | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı ve Öğrenci | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.title | Pre-Trained Variational Autoencoder Approaches for Generating 3D Objects from 2D Images | |
| dc.type | Conference Object |












