Vol. 1, Issue 3 β’ 2026
Deepanshi Joon, Aditi Nautiyal
Unlike stone age, humans no longer live in caves or mud houses but rather shifted to modern civilization setups which include - buildings, bridges and industrial facilities as their environment. And with time these structures and components develop surface defects like cracks, peeling, spalling, algae, staining and manufacturing flaws. If frequently, these structures are left untreated, then it may affect human safety and threaten serviceability. For inspection purposes, only relying on manual visual examination will lead to a slow process and some parts may be left unchecked. This will add pressure on pockets. Moreover, this method is slow, subjective, costly, error-prone and unsafe for surfaces which are inaccessible. Overall, these factors motivate automated computer vision inspection. However, there exists some constraints like limited, imbalanced datasets and inspection spans of various distinct types of tasks. This work gives unified Deep Learning model which evaluates on one of the public benchmark datasets from Kaggle which consists of sub folders of Magnetic-Tile Defect, DeepPCB, DBCC bridge cracks and CrackForest. In this paper multi-class classification, binary classification and pixel-level segmentation are considered. The model coupleβs residual backbone with classification, which is regularized head and symmetric U-Net. Overall is trained from scratch without pre-training and employs objective which is imbalanced (class-weighted, positive weighted BCE with Dice, label-smoothed cross entropy). This framework attains F1 score of 0.9945 for DeepPCB dataset. The novelty lies in its approach, as its single, pre-training free pipeline which spans three task paradigms with the issue of handling imbalance and format agnostic ingestion and in real honest finding that scale of data, more than architecture and govern effectiveness. Future work could focus on adding pre- trained backbones, multi-seed evaluation and board- disjoint and then end to end detection.
Palvi Sharma
Multi-crop plant disease automatic classification in precision agriculture requires an effective solution but the state-of-the-art deep learning architectures are prone to suffering background shortcut problem, higher inference latency, and inadequate visualization interpretability. This paper presents a better performing framework combining ConvNeXt-Tiny with an innovative Residual Spatial Attention Module (RSAM) to solve the fine-grained diagnosis problem for 38 plant pathology targets. On a dataset of 10,876 unseen images, the proposed framework demonstrates a Top-1 classification accuracy of 96.85%, a precision of 96.95%, a recall of 96.48%, and a macro F1-score of 96.71% while being superior to ResNet- 50 (+2.20%) and being 0.97 ms faster per image (7.15 ms/img with 28.12 M parameters). Explainable AI (XAI) analysis with the help of Grad-CAM shows the ability of the RSAM block to suppress background and soil artifacts, directing 88.42% of activation energy πΈπππ πππ
Shubham Gupta
Railway compressor systems are essential to the pneumatic functions that keep metro operations safe and uninterrupted, yet unexpected degradation of these systems remains a persistent source of service disruption and unplanned maintenance costs. This study proposes an Explainable Temporal Data Analytics (ETDA) framework for early failure prediction and anomaly detection in railway compressor systems, using the publicly available MetroPT-3 dataset [1,2], which contains 15,169,480 one-second observations from 15 analogue and digital sensors of a metro Air Production Unit collected between February and August 2020. Rather than treating sensor readings as independent samples, the proposed framework transforms instantaneous measurements into rolling temporal features (mean, standard deviation, rate of change, range, and cross-sensor correlation), derives a persistence-aware anomaly score, and estimates multi-horizon failure risk, coupled with Shapley-based feature attribution to explain which behavioural changes drive the estimated risk. Chronological validation is used throughout to avoid temporal leakage. The framework was evaluated against five conventional classifiers and detected all four documented air-leak failures, with a mean anomaly lead time of 6.8 hours. At the practically useful 6-hour warning horizon, the framework achieved a 91.4% F1-score and a PR-AUC of 0.903, exceeding the strongest conventional baseline (CatBoost, 88.3% F1-score) by 3.1 percentage points. Ablation analysis confirmed that temporal features, anomaly persistence, and the composite risk-integration layer each contribute statistically significant improvements (Wilcoxon signed-rank test, p < 0.05). Explainability analysis identified motor-current and pressure-related rolling features as the dominant drivers of predicted failure risk, results that are consistent with the physical behaviour of the compressor. The results presented here show that the integration of temporal analytics, persistent anomaly detection, and explainable early-warning estimation has the potential to furnish a more informative, actionable, and helpful foundation for railway maintenance decision-making than approaches based on instantaneous sensor measurements.
Shivareddy Devarapalli, Monu Sharma, Abhishek Jain
Enterprise resource planning (ERP) systems that are based on workdays are becoming dependent on extensive, heterogeneous and rapidly changing financial information, but current tools are constrained by restricted data structures, hand written rules, and poor generalization among enterprises. When implemented in the context of organizations of varying transaction volumes, document structures, and risk indicators, these restrictions pose serious impediments to the automation of major financial processes, namely, wire transfers, payment to employees and budget approvals. To fill this gap, a scalable cross-enterprise data modelling is suggested that can learn unified features using heterogeneous financial logs and enhance adaptive automation when applying Workday ERP processes. The proposed methodology utilizes a differentiable deep neural network (DDNN) predictive model, anomaly detection and workflow risk scoring; and a modified dung beetle optimization (MDBO) algorithm is used to optimize the routing of decisions, reduce processing delays and improve the selection of multi-domain features. An inter-enterprise benchmark dataset is built by aligning structured and unstructured financial data with the help of ontology-based normalization, and the DDNN-MDBO model is tested using two example ERP processes inter-bank wire transfer and employee reimbursement cycles. The comparison of two workflows indicates that the DDNN of 97.1% and MDBO model have 98.1% and 96.2% accuracy on inter-bank transfers and reimbursements, respectively, with the precision, recall and ROC AUC scores being also high and decreasing false positives/negatives by 42%. MDBO increases risk control and parallelization by 100 percent and cuts the processing time by 28% shortens turnaround time of cross enterprise ERP automation is robust and scalable, and predictive performance and operational efficiencies are improved.
Mohammad Shahnawaz Shaikh
Text to video generation has advanced significantly in recent years, largely due to the development of extremely sophisticated diffusion models. In this work, we present a novel ap- proach to producing excellent video content based on descriptions by utilizing diffusion tech- niques. Using a multi-stage diffusion process, we describe a framework that progressively trans- forms text inputs into coherent video sequences. To address the intricate spatial and temporal idiosyncrasies of video data, we construct a model that combines a robust text encoder with a diffusion network that has been specifically optimized. To ensure that the movies generated are both visually appealing and contextually sound, our approach combines attention mechanisms with temporal consistency restrictions. We test our approach on several datasets and show that it significantly outperforms current methods in terms of video quality, relevance to textual prompts, and temporal continuity. Importantly, our findings highlight the usefulness of diffusion models for text-to-video production, showing a viable path toward producing dynamic and engaging content from textual descriptions. In addition to expanding the potential of generative models, this study establishes the framework for innovative new uses in a variety of industries, including education and entertainment. Future studies should aim to scale the model and expand the variety of video outputs in terms of resolution. The application of generative models in diffusion models has been incredibly successful, and the area has demonstrated amazing promise in a variety of fields, including the creation of images and sounds.
PRASHANT SONI, Yuvank Soni
Web3 apps offer decentralized and user-centric digital services through blockchain networks, smart contracts, and pseudonymous wallet identities. Nevertheless, the transparency and evolving nature of Web3 environments create security issues such as illicit transactions, credentials theft, wallet anomalies, and access risks based on contextual factors. Traditional access control methods typically depend on static identity, predetermined permissions, or fixed rules so they can be inefficient when the security status of a user or transaction changes instantly. This paper suggests a context-aware risk scoring and adaptive access control system to be used in secure Web3 applications. The proposed system relies on a range of contextual factors to calculate a dynamic risk score for each of the access requests or transactions. Risky requests are allowed, requests with medium-level risks are subjected to extra verification needs, while the requests categorized under high-level risk are turned down. For the model to recognize behavioral changes and security condition shifts, it uses continuous contextual evaluation in the middle of active sessions. The aim of the experimental evaluation is to assess the capability of the system to detect attacks and determine its rate of false positive results, the effectiveness of its classification, the latency of its decision-making, and the improvement of security.