Saas and Web design
A government-backed evaluation framework for LLM applications ensuring safer AI deployments in India
Year :
2025
Tools :
Figma, Miro, Bolt AI,
Claude AI, Chat GPT
Responsibilities
Research:
Amrita University


Outcome: Designed 20+ screens and 6 dashboards, translating complex AI test results into a 1- 100 risk scoring system through 8 feedback cycles with 6- 7 stakeholders.
Learning: In high-stakes AI systems, clear thresholds, hierarchy, and accessibility are essential for trustworthy decisions.
The Challenge
India faces critical AI safety risks across three vital sectors:
Healthcare: Unsafe LLM medical advice risks patient safety
Finance: Misleading AI guidance leads to exploitation of citizens
Education: Incorrect content creates learning inequities
Problem
While global governments create AI safety frameworks, existing solutions are generic and fail to capture India's diversity of languages, cultures, and socio-economic need
Opportunity
Track LLM is a context-aware evaluation framework that enables developers, testers, and GRC teams to evaluate their LLM applications for bias, fairness, toxicity, and truthfulness, specifically within India's diverse context.
Discovery & Research Phase
Problem Space
Global AI Safety Context:
Governments worldwide creating AI safety frameworks
Focus on risk assessments, bias checks, accountability
Generic solutions don't address India's unique context
India-Specific Challenges Identified:
Linguistic Diversity: 22+ scheduled languages underrepresented in LLMs
Cultural Context: Caste, religion, regional biases not addressed
Socio-Economic Disparity: Rural vs urban, economic class variations
Domain Criticality: Healthcare, Finance, Education require highest safety
Key Insight- Existing LLM evaluation tools test models (e.g., 'Is GPT-5 biased?') but don't evaluate real-world applications (e.g., 'Is my Tutoring Bot safe for Indian students?')
Competitor Analysis
Tools Analyzed:
HELM, Latitude, LangWatch, LM Eval Harness
OpenAI Evals, PromptBench, MT-Bench
Competitive Advantage Identified:
First platform to evaluate LLM applications vs. models
Domain-specific risk assessment for India
Actionable analytics that show how to fix problems
Key Research Findings
Pain Points Discovered:
Developers:
"We don't know if our chatbot is safe to deploy"
"Generic bias tests don't catch India-specific issues"
"We need domain-specific evaluation for healthcare apps"
"Current tools are too technical for our team"
GRC Teams:
"Need compliance documentation for AI governance"
"Can't explain AI risks to non-technical stakeholders"
"Lack standardized evaluation frameworks for India"
Researchers:
"No datasets covering Indian cultural contexts"
"Western benchmarks miss caste, regional biases"
"Need reproducible evaluation methodology"
Information Architecture

Low-fi designs
Design System
Track LLM Color System
A modern, accessible color palette for India's AI evaluation platform
Primary Colors
Primary Lavender
#A7ABF6
Main brand color for headers, navigation, primary actions, and key UI elements.
Represents innovation and approachability.
✓ WCAG AA
Accent Olive
#819337
Main brand color for headers, navigation, success states, positive metrics. Represents growth, balance, and natural intelligence.
✓ WCAG AA
Color Psychology
Lavender: Calming yet modern. Breaks from traditional government blues while maintaining professionalism. Creates approachable AI platform feel
without appearing frivolous.
Olive Green: Grounded and trustworthy. Natural green holds universally positive connotations in Indian culture. Balances tech-forward lavender with organic stability.
Semantic Colors
Success Green
7FD798
Grade A, excellent performance,
completed states, positive outcomes.
Warning
#FFD19C
Grade B/C, moderate risk, areas
needing attention and review.
Critical Red
#F55E5E
Grade D/F, critical issues, errors
requiring immediate attention.
Information Blue
#7C7EEB
Informational alerts, neutral highlights.
Derived from primary lavender.
Neutral Palette
Gray 50
#F8F7F7
Page backgrounds, light surfaces
Gray 100
#F3F4F6
Card backgrounds, hover states
Gray 300
#D1D5DB
Borders, dividers
Gray 500
#6B7280
Secondary text, icons
Gray 900
#262626
Primary text, headings
Design Rationale
Why this palette works for Track LLM:
Modern yet Trustworthy: Lavender breaks from traditional government blues while maintaining institutional credibility
Balanced Contrast: The cool lavender + warm olive creates visual interest without overwhelming
Accessible: All color combinations meet WCAG AA standards for readability
AI-Forward: Purple/lavender hues are associated with innovation and technology without being cliché
Calming Interface: Softer colors reduce anxiety when dealing with AI safety issues
Usage Distribution:
60% Neutral grays (structure, backgrounds)
25% Primary Lavender (navigation, headers, primary actions)
10% Accent Olive (CTAs, positive highlights)
5% Semantic colors (alerts, grades)
Outcomes
Delivered 20+ screens, including 6 data-heavy dashboards, for a government-backed LLM benchmarking platform.
Designed 1 primary dashboard + 5 analytical views to visualise AI evaluation results using a 1- 100 risk scoring system.
Implemented contextual risk alerts to highlight high-severity and non-compliant model behaviour.
Iterated across 8 major feedback cycles over 8 months, collaborating with 6-7 stakeholders across AI/ML, product, engineering, and leadership.
Applied accessibility best practices to dashboards, including colour-contrast validation, to support inclusive and responsible interpretation of AI risk data.
Learnings
Numeric scoring systems (1- 100) require clear thresholds and hierarchy to avoid misinterpretation.
Data-dense dashboards need progressive disclosure to balance overview and depth.
Long-running projects benefit from structured feedback cycles to manage stakeholder complexity.
Accessibility is critical in risk-based interfaces, where clarity directly impacts decision-making.












