SmartCloud: Extending an AI-Powered Cloud Storage System with Automated S3 Glacier Archiving, SimHash Near-Duplicate Detection, and a Custom-Trained Random Forest Recommendation Engine

Year : 2026 | Volume : 13 | Issue : 02 | Page : 36 44
By

Rutuja K. Taware,

Vishal Shep,

Kartik Sable,

Yash Tilekar,

Ankit Bhujbal,

  1. Assistant Professor, Department of Computer Engineering, SVPM’s College of Engineering, Maharashtra, India
  2. Student, Department of Computer Engineering, SVPM’s College of Engineering, Maharashtra, India
  3. Student, Department of Computer Engineering, SVPM’s College of Engineering, Maharashtra, India
  4. Student, Department of Computer Engineering, SVPM’s College of Engineering, Maharashtra, India
  5. Student, Department of Computer Engineering, SVPM’s College of Engineering, Maharashtra, India

Abstract

Cloud storage services have become fundamental infrastructure for both individuals and organizations, yet they continue to charge users for every byte retained without offering any mechanism to automatically reduce that footprint over time. Files accumulate progressively, near-identical document revisions go undetected, and cold data occupies expensive active storage tiers indefinitely. The present work introduces SmartCloud, a fully deployed multi-cloud storage optimization platform that addresses these inefficiencies through four integrated engineering mechanisms. First, a nightly archiving pipeline built on a node-cron scheduler evaluates every active, non-pinned file against a 30-day inactivity threshold and a 50 MB size threshold, migrating qualifying files to AWS S3 Glacier in an atomic three-step operation and achieving a per-gigabyte cost reduction of 82.6 percent compared to S3 Standard pricing. Second, a dual-layer deduplication engine combines SHA-256 exact hash matching with SimHash 64-bit near-duplicate fingerprinting, where each file buffer is divided into fixed 4096-byte chunks, hashed with MD5, and reduced to a 64-bit weighted fingerprint through a bit-voting process. Files whose fingerprints fall within a Hamming distance of 10 bits from any stored record are flagged as near-duplicates, detecting modified file copies that byte-level comparison cannot identify. Third, a Random Forest classifier trained on 2,000 synthetically generated storage profiles across five recommendation classes is served through a FastAPI microservice with inference latency below 10 milliseconds, providing data-driven storage optimization recommendations on every dashboard session. Fourth, per-file pin and restore controls allow users to exempt files from automated archiving and retrieve archived content from Glacier on demand. The system is deployed across React on Vercel, Node.js on Render, Supabase, AWS S3 and Glacier, and MongoDB Atlas. Evaluation estimates 40 to 60 percent total active storage reduction on a representative 100 GB mixed-format dataset.

Keywords: S3 Glacier Archiving, SimHash Deduplication, Random Forest, Cloud Storage Optimization, FastAPI Microservice, Brotli Compression, Multi-Cloud, BERT Semantic Search

[This article belongs to Journal of Software Engineering Tools & Technology Trends ]

How to cite this article: Rutuja K. Taware, Vishal Shep, Kartik Sable, Yash Tilekar, Ankit Bhujbal. SmartCloud: Extending an AI-Powered Cloud Storage System with Automated S3 Glacier Archiving, SimHash Near-Duplicate Detection, and a Custom-Trained Random Forest Recommendation Engine. Journal of Software Engineering Tools & Technology Trends. 2026; 13(02):36-44.
How to cite this URL: Rutuja K. Taware, Vishal Shep, Kartik Sable, Yash Tilekar, Ankit Bhujbal. SmartCloud: Extending an AI-Powered Cloud Storage System with Automated S3 Glacier Archiving, SimHash Near-Duplicate Detection, and a Custom-Trained Random Forest Recommendation Engine. Journal of Software Engineering Tools & Technology Trends. 2026; 13(02):36-44. Available from: https://journals.stmjournals.com/josettt/article=2026/view=259421

References

  1. Amazon Web Services. How it works—S3 Intelligent-Tiering [Internet]. 2026. Available from: https://aws.amazon.com/s3/storage-classes/intelligent-tiering/
  2. Nimmagadda S. Intelligent data tiering with access-aware storage protocols: architectures, algorithms, and applications. Int J Innov Res Sci Eng Technol. 2023;12.
  3. Kanthavel R, Dhaya R, Khalaf OI, Hamad AA. Artificial intelligence-based smart cloud computing schema model. Control Cybern. 2024;53:639-72. doi:10.2478/candc-2024-0025.
  4. Balamanigandan R, Mahaveerakannan R, Saraswathi S, Jenifer AM. Optimizing resource costs with machine learning techniques. In: Proceedings of the International Conference on Machine Learning and Autonomous Systems (ICMLAS); 2025. p. 364-70. doi:10.1109/ICMLAS64557.2025.10968124.
  5. Gautam D, Saxena V. Optimization of storage of cloud servers through binary search algorithm. In: Proceedings of the IEEE 7th Conference on Information and Communication Technology (CICT); 2023. p. 1-6. doi:10.1109/CICT59886.2023.10455361.
  6. Yen FI, Rahman MM. Efficient image compression for cloud system. In: Proceedings of the International Conference on Sustainable Technologies for Industry 4.0 (STI); 2019. p. 1-6. doi:10.1109/STI47673.2019.9067989.
  7. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Vol. 1, Long and Short Papers; 2019. p. 4171-86. doi:10.18653/v1/N19-1423.
  8. Alakuijala J, Farruggia A, Ferragina P, Kliuchnikov E, Obryk R, Szabadka Z, et al. Brotli: a general-purpose data compressor. ACM Trans Inf Syst. 2019;37:1-30. doi:10.1145/3231935.
  9. Anand K. AI-driven optimization of cloud data centers: enhancing reliability and cost efficiency. In: Proceedings of the 2025 International Conference on Data Science, Agents & Artificial Intelligence (ICDSAAI); 2025. p. 1-6.
  10. Breiman L. Random forests. Mach Learn. 2001;45:5-32. doi:10.1023/A:1010933404324.
  11. Zhu B, Li K, Patterson RH. Avoiding the disk bottleneck in the data domain deduplication file system. In: Proceedings of FAST. 2008.
  12. Charikar MS. Similarity estimation techniques from rounding algorithms. In: Proceedings of the Thirty-Fourth Annual ACM Symposium on Theory of Computing; 2002. p. 380-8. doi:10.1145/509907.509965.
  13. Pokharana A, Sharma S. Encryption, file splitting and file compression techniques for data security in virtualized environment. In: Proceedings of the Third International Conference on Inventive Research in Computing Applications (ICIRCA); 2021. p. 480-5. doi:10.1109/ICIRCA51532.2021.9544599.

 


Regular Issue Subscription Original Research
Volume 13
Issue 02
Received 05/05/2026
Accepted 10/07/2026
Published 25/08/2026
Publication Time 112 Days


Login

My IP

PlumX Metrics

Support