Hyderabad, Telangana, India
3K followers 500+ connections

Join to view profile

About

I'm currently a Principal Scientist & Director of Applied AI at NetApp focusing on…

Activity

3K followers

See all activities

Experience & Education

  • NetApp

View Phani’s full experience

See their title, tenure and more.

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

Licenses & Certifications

Volunteer Experience

  • Volunteer

    Microsoft

    - 1 year 1 month

    Social Services

    Giving Activities Coordinator for my data science & research organization.

Publications

  • Open Your Eyes: Benchmarking the Detection of Fabricated Realities and Weaponized Ethics in VLMs

    ECCV

    As Vision–Language Models (VLMs) increasingly power autonomous agents—from coding assistants that interpret terminal output and execute shell commands, to Computer-Use Agents with full desktop control, to HR screening and loan evaluation bots—these agents act directly over visual artifacts from untrusted sources. While prior work has studied direct multimodal jailbreaks, text-based indirect prompt injection, and HTML-embedded attacks, the security implications of Multimodal Indirect Prompt…

    As Vision–Language Models (VLMs) increasingly power autonomous agents—from coding assistants that interpret terminal output and execute shell commands, to Computer-Use Agents with full desktop control, to HR screening and loan evaluation bots—these agents act directly over visual artifacts from untrusted sources. While prior work has studied direct multimodal jailbreaks, text-based indirect prompt injection, and HTML-embedded attacks, the security implications of Multimodal Indirect Prompt Injection (M-IPI) in realistic agentic workflows remain insufficiently understood. We introduce the M-IPI Detection Benchmark, a suite of 2,600 high-fidelity visual artifacts encompassing two attack families: (1) Technically-Framed Attacks, where malicious commands are interwoven with genuine debugging workflows in terminal screenshots, and (2) Ethics-Framed Attacks, a novel vector where adversaries exploit alignment priors such as fairness mandates to override evaluation policies. We evaluate five model families in paired vision and text-only configurations that receive identical content, isolating the modality effect. Our results reveal that VLMs consistently underperform text-only counterparts at detecting embedded attacks—a "visual authority" effect where rendered presentation suppresses critical evaluation. For end-to-end attack success, both modalities show comparable susceptibility (~50% for technical, >80% for ethics-framed), with qualitatively distinct failure profiles: VL models disproportionately trust visually-rendered commands while text-only models are more vulnerable to natural-language directives. These findings indicate that the vulnerabilities are rooted in the shared language backbone and that alignment priors represent a distinct, exploitable attack surface.

    Other authors
    See publication

Patents

  • Data Exfiltration Monitoring Using Semantic Queries

    Filed US-20250330487-A1

    The disclosure describes a data protection service that generates semantic descriptions of protected data volumes. The data protection service queries a monitoring service with the generated semantic descriptions. The monitoring service responds to the queries with indications of whether and data items on the dark web match the semantic descriptions. When a query receives a positive response from the monitoring service, the data protection service iteratively refines the semantic description…

    The disclosure describes a data protection service that generates semantic descriptions of protected data volumes. The data protection service queries a monitoring service with the generated semantic descriptions. The monitoring service responds to the queries with indications of whether and data items on the dark web match the semantic descriptions. When a query receives a positive response from the monitoring service, the data protection service iteratively refines the semantic description and queries the monitoring service with the refined semantic descriptions until a breach is detected. Once a breach is detected, the data protection service initiates a mitigation action.

    See patent
  • Data Preparation Engine(s) For Curating Secure And Compliant Data Collections From Distributed Sources

    Filed US-20260080099-A1

    Various embodiments of the present technology generally relate to systems and methods for providing a data preparation engine for curating secure and compliant data collections from distributed storage systems. In an aspect, a data preparation engine receives a query from a client device and determines files from one or more distributed sources based on the query. The data preparation engine determines sensitive data within the files and anonymizes the sensitive data while preserving context…

    Various embodiments of the present technology generally relate to systems and methods for providing a data preparation engine for curating secure and compliant data collections from distributed storage systems. In an aspect, a data preparation engine receives a query from a client device and determines files from one or more distributed sources based on the query. The data preparation engine determines sensitive data within the files and anonymizes the sensitive data while preserving context and integrity of the underlying information. The data preparation engine generates a data collection including the files with anonymized sensitive data. The data collection may then be deployed to downstream applications or workflows, such as used to generate curated data sets for training of artificial intelligence applications. Once deployed, the data preparation engine may continuously monitor the distributed sources for changes to data within the files and automatically update data collections in real-time.

    See patent
  • Graph Vector Variation Driven Data Corruption Detection

    Filed US-20250245326-A1

    Disclosed herein are methods, systems, and apparatus for the detection of data integrity anomalies indicative of malware for a datastore of an organization. To identify an anomaly in a file, a portion of a file is identified to be used in a vector comparison. The portion can comprise sentences or paragraphs for text files, entries, rows, or columns for spreadsheet files, or some other divisible portion of a file. A vector having multiple dimensions is generated for the portion based on the…

    Disclosed herein are methods, systems, and apparatus for the detection of data integrity anomalies indicative of malware for a datastore of an organization. To identify an anomaly in a file, a portion of a file is identified to be used in a vector comparison. The portion can comprise sentences or paragraphs for text files, entries, rows, or columns for spreadsheet files, or some other divisible portion of a file. A vector having multiple dimensions is generated for the portion based on the content in the portion. Each dimension of the multiple dimensions corresponds to a feature of the portion. A variation is determined between the vector and one other vector associated with one other portion of the file. One or more actions to take with respect to the file is determined based on the variation, such as malware mitigation, and the action is performed with respect to the file.

    See patent
  • Ransomware Detecting using Decoy Files

    Filed US-20250328644-A1

    Disclosed herein are systems, methods, and software for the operation of a ransomware detection system. The ransomware detection system generates a decoy file based on characteristics of an existing file in a file system. The decoy file is effectively indistinguishable from the existing file from the perspective of the ransomware but contains simulated data rather than authentic data. The ransomware detection system identifies a location in the file system and deploys the decoy file to the…

    Disclosed herein are systems, methods, and software for the operation of a ransomware detection system. The ransomware detection system generates a decoy file based on characteristics of an existing file in a file system. The decoy file is effectively indistinguishable from the existing file from the perspective of the ransomware but contains simulated data rather than authentic data. The ransomware detection system identifies a location in the file system and deploys the decoy file to the location. The decoy is then monitored to detect changes by comparing a ground truth for the decoy file to the current state of the decoy file. The decoy file is checked for changes at a rate associated with the identified location. Where a change is detected, an alert is sent to a ransomware mitigation process, which initiates ransomware mitigation.

    See patent
  • Data Exfiltration Monitoring Using Hash Values

    Filed US-20250330488-A1

    The disclosure describes a data protection service that generates semantic descriptions of protected data volumes. The data protection service queries a monitoring service with the generated semantic descriptions. The monitoring service responds to the queries with indications of whether and data items on the dark web match the semantic descriptions. When a query receives a positive response from the monitoring service, the data protection service iteratively refines the semantic description…

    The disclosure describes a data protection service that generates semantic descriptions of protected data volumes. The data protection service queries a monitoring service with the generated semantic descriptions. The monitoring service responds to the queries with indications of whether and data items on the dark web match the semantic descriptions. When a query receives a positive response from the monitoring service, the data protection service iteratively refines the semantic description and queries the monitoring service with the refined semantic descriptions until a breach is detected. Once a breach is detected, the data protection service initiates a mitigation action.

    See patent

Honors & Awards

  • 2nd prize - H2o.ai Predict the LLM Kaggle Hackathon

    Kaggle

    The objective of this NLP competition is to detect which out of 7 possible LLM models produced a particular output. With each model having its unique subtleties and quirks, participants must build accurate models to identify the source of an answer. The challenge of pinpointing the origin LLM of a given output is not only intriguing but is also an area of spirited research.

    Out of 106 participants, I stood 2nd in this text classification contest.

    Leaderboard:…

    The objective of this NLP competition is to detect which out of 7 possible LLM models produced a particular output. With each model having its unique subtleties and quirks, participants must build accurate models to identify the source of an answer. The challenge of pinpointing the origin LLM of a given output is not only intriguing but is also an area of spirited research.

    Out of 106 participants, I stood 2nd in this text classification contest.

    Leaderboard: https://www.kaggle.com/competitions/h2oai-predict-the-llm/leaderboard
    Solution: https://www.kaggle.com/code/phanisrikanth/2nd-place-solution-decoder-inference
    ML Logbook: https://github.com/binga/kaggle-predict-the-llm/blob/main/ml_logbook.md

  • Winner - Identifying Superheroes from Product Images

    CrowdAnalytix

    The dataset in this challenge consisted of product images like t-shirts, bags etc. with superhero graphics. In this contest, the participants are tasked to build machine learning models to identify the superheroes in an image (fashion product images).

    Leaderboard: https://www.crowdanalytix.com/contests/identifying-superheroes-from-product-images

  • 2nd runner-up - Recommender Systems Machine Learning Challenge

    HackerEarth

    Hotstar is an on demand video streaming service in India. It boasts about having more than a 100 Million users and more than 35,000 hours of content on their platform. Due to the sheer scale of users and content, leveraging user browsing history and thus personalising the content for each user creates a lot of value for Hotstar.
    In this challenge, Hotstar challenged us with building a recommendation system so that they could personalize the user experience and also improve the content…

    Hotstar is an on demand video streaming service in India. It boasts about having more than a 100 Million users and more than 35,000 hours of content on their platform. Due to the sheer scale of users and content, leveraging user browsing history and thus personalising the content for each user creates a lot of value for Hotstar.
    In this challenge, Hotstar challenged us with building a recommendation system so that they could personalize the user experience and also improve the content consumption on their platform (and hence revenues!).

    Techniques: Collaborative Filtering, Deep Learning using Keras, TensorFlow, Python.
    Blog Post: https://medium.com/data-science-analytics/building-a-movie-recommendation-engine-for-hotstar-478fb4b21c17

  • Winner - Analytics Roadshow MiniHack

    Analytics Vidhya

    The dataset provided belongs to Sigma Cab Private Limited - a cab aggregator service. Their customers can download their app on smartphones and book a cab from any where in the cities they operate in. They, in turn search for cabs from various service providers and provide the best option to their client across available options. They have been in operation for little less than a year now. During this period, they have captured surge_pricing_type from the service providers.

    The objective…

    The dataset provided belongs to Sigma Cab Private Limited - a cab aggregator service. Their customers can download their app on smartphones and book a cab from any where in the cities they operate in. They, in turn search for cabs from various service providers and provide the best option to their client across available options. They have been in operation for little less than a year now. During this period, they have captured surge_pricing_type from the service providers.

    The objective of this challenge is to build a predictive model, which could help them in predicting the surge_pricing_type pro-actively. This would in turn help them in matching the right cabs with the right customers quickly and efficiently.

    Techniques: Gradient Boosted Trees using XGBoost, Python.
    Reference: https://datahack.analyticsvidhya.com/contest/minihack-machine-learning/lb
    Blog post: https://medium.com/data-science-analytics/winning-two-machine-learning-challenges-in-the-same-month-6f36d0ca28a6

  • Winner - Recommender Systems Machine Learning Challenge

    Analytics Vidhya

    Understanding customers and their preferences is the holy grail for online businesses. Building a recommender system is one of the common ways to do so.

    The objective of this contest is to build a model that predicts a given user’s ratings (from 0 to 10 stars) for a given item based on past ratings on other items and/or other information. No additional information (user demographics, item content features etc.) are given and the prediction has to be made using only the historical ratings…

    Understanding customers and their preferences is the holy grail for online businesses. Building a recommender system is one of the common ways to do so.

    The objective of this contest is to build a model that predicts a given user’s ratings (from 0 to 10 stars) for a given item based on past ratings on other items and/or other information. No additional information (user demographics, item content features etc.) are given and the prediction has to be made using only the historical ratings of items.

    Techniques: Collaborative Filtering, Matrix Factorization, Deep Learning, Gradient Boosted Trees using Keras, TensorFlow, LightGBM, Python.
    Reference: https://datahack.analyticsvidhya.com/contest/mlware-2/lb
    Blog post: https://medium.com/data-science-analytics/winning-two-machine-learning-challenges-in-the-same-month-6f36d0ca28a6

  • Winner - Fraud Detection Machine Learning Challenge

    HackerEarth and Societe Generale

    Societe Generale, one of the largest banks in France, in collaboration with HackerEarth, organised Brainwaves, the annual hackathon at Bengaluru on November 12–13, 2016. The theme of the hackathon this year was “Machine Learning”. The hackathon had an online qualifier from where 85 top teams out of 2200 registrations from all over India, were selected for the final round. We finished 1st amongst all the teams.

    Techniques: Gradient Boosted Trees, Random Forests using XGBoost…

    Societe Generale, one of the largest banks in France, in collaboration with HackerEarth, organised Brainwaves, the annual hackathon at Bengaluru on November 12–13, 2016. The theme of the hackathon this year was “Machine Learning”. The hackathon had an online qualifier from where 85 top teams out of 2200 registrations from all over India, were selected for the final round. We finished 1st amongst all the teams.

    Techniques: Gradient Boosted Trees, Random Forests using XGBoost, Scikit-Learn, Python.
    Reference: https://www.hackerearth.com/brainwaves/
    Blog Post: https://medium.com/data-science-analytics/winning-the-hackerearth-machine-learning-challenge-2038cc72401c#.3xa7qzav0

  • Winner - Text Mining Machine Learning Challenge

    CrowdAnalytix

    Millions of dollars are spent developing software that maintain information about products, buying history of users for particular products, etc. But as the catalogue size and no. of suppliers keeps growing the problem of maintaining this catalogue accurately grows exponentially. One of the attributes which is important for e-tailers is MPN (Manufacturer Part Number). MPN is unique identifier assigned to a product by the manufacturer. MPN, generally present as a part of Title/Description, helps…

    Millions of dollars are spent developing software that maintain information about products, buying history of users for particular products, etc. But as the catalogue size and no. of suppliers keeps growing the problem of maintaining this catalogue accurately grows exponentially. One of the attributes which is important for e-tailers is MPN (Manufacturer Part Number). MPN is unique identifier assigned to a product by the manufacturer. MPN, generally present as a part of Title/Description, helps the buyers to check for authenticity of the product. The objective of the challenge is to extract the MPN for a given product from its Title/Description using regular expressions.

    Techniques: Random Forests, Text Analytics, Regex using Scikit-Learn, Python.
    Reference: https://www.crowdanalytix.com/contests/extraction-of-product-attribute-values

  • Winner - Predict Customer Worth For Banks

    Analytics Vidhya

    Digital arms of banks today face challenges with lead conversion, they source leads through mediums like search, display, email campaigns and via affiliate partners. The objective of this contest is to identify the customers segments having higher conversion ratio for a specific loan product so that they can specifically target these customers.

    Blog post: https://medium.com/data-science-analytics/analytics-vidhya-3-x-hackathon-9f2550b47be6

  • 2nd prize - Marketing Analytics Hackathon - DataMeet Mumbai

    http://www.meetup.com/DataMeet-Mumbai/events/223625039/

    Propensity model development –the client has a couple of use cases where they have not been able to get 80% response capture in top 3 deciles / >3X lift in the top decile - inspite of several iterations. The expectation here would be identification of any new technique / algorithm (apart from logistic regression), which can help get the desired model performance.

    Blog post: https://medium.com/data-science-analytics/data-science-hackathon-datameet-mumbai-bbf080e3009b

  • Gold Medal for Excellence in Research

    Department of Electrical Engineering, NIT Warangal

    I was awarded a Gold Medal from Head of Department for my research work that got accepted in International Machine Learning conferences and was published in journals.

  • Merit Scholarship Awardee

    -

    Awarded merit scholarship for outstanding performance in engineering entrance examinations.

Languages

  • English

    Native or bilingual proficiency

  • Hindi

    Professional working proficiency

  • Telugu

    Native or bilingual proficiency

  • Python

    Full professional proficiency

View Phani’s full profile

  • See who you know in common
  • Get introduced
  • Contact Phani directly
Join to view full profile

Other similar profiles

Explore collaborative articles

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Explore More

Others named Phani Srikanth

Add new skills with these courses