Your Articles, Anywhere · 2025-12-07 · 43 min
Key moments - from our scoring
Substance score
56 / 100
Five dimensions, 20 points each
Clio addresses a critical gap in AI governance: the lack of empirical data about how AI assistants are actually used at scale. While model providers have access to millions of conversations, privacy concerns, ethical issues around human review, competitive pressures, and sheer volume have prevented systematic analysis. The platform uses AI assistants themselves to extract aggregated insights - similar to how Google Trends anonymizes search behavior - transforming raw conversations through multi-stage pipelines that extract topics, cluster semantically similar content, and generate privacy-preserving descriptions. Applied to 1 million Claude conversations, Clio reveals that coding, writing, and research tasks dominate usage with significant cross-linguistic variations (e.g., Japanese conversations discuss elder care at higher rates). Beyond usage patterns, Clio serves critical safety functions: identifying coordinated abuse attempts, monitoring risks during capability launches and elections, and improving existing safety classifiers. The platform demonstrates that empirical observation of real-world AI usage can inform governance frameworks, product decisions, and proactive safeguards without compromising user privacy or requiring human reviewers to read sensitive data.
Clio is a privacy-preserving platform that uses AI assistants to analyze millions of conversations by extracting aggregated patterns through a multi-stage pipeline - extracting facets like topic and language, clustering semantically similar conversations, generating privacy-preserving descriptions, and organizing them into navigable hierarchies. It achieves 94% accuracy in reconstructing topic distributions while maintaining undetectable private information leakage through multiple layers of statistical safeguards.
Coding-related tasks dominate usage at over 10% of conversations (especially web and mobile development), followed by writing and communication tasks (professional email, document creation), and research/educational uses (6-10% of conversations). Cross-linguistic analysis shows Japanese conversations discuss elder care at higher rates while Spanish conversations feature more economic theory discussions.
Clio identifies coordinated abuse attempts, monitors for unknown risks during high-stakes periods like capability launches and elections, and improves existing safety classifiers by surfacing patterns invisible at the individual conversation level. Its ability to cluster similar conversations across millions of interactions reveals sophisticated attacks like systematic SEO spam generation.
Four fundamental challenges prevent sharing: privacy concerns from users sharing sensitive personal and business information; ethical issues with human reviewers examining potentially distressing content; competitive pressures to protect user base intelligence from rivals; and practical infeasibility of manual review at scale with millions of daily messages.
Organizations should prioritize transparent communication about capabilities and limitations, implement procedural justice in content moderation (explanation, appeals, human oversight), invest in user capability building through domain-specific guidance and critical evaluation training, and provide targeted access and financial support for beneficial use cases like education and safety research.
Our reviewer’s read on each dimension, with quotes from the episode.
The transcript delivers substantial, non-obvious insights about real-world AI usage patterns grounded in analysis of 1 million Claude conversations. Novel findings include cross-cultural usage variations (e.g., Japanese elder care discussions), the dominance of coding tasks (10%+ of conversations), and coordinated abuse detection methods. However, significant portions are devoted to restating organizational philosophy and governance principles that, while important, dilute insight density toward the latter half.
coding related tasks dominate usage, with web and mobile application development representing over 10% of all conversations
Japanese and Chinese conversations show elevated rates of elder care discussions compared to other languages, potentially reflecting demographic challenges these societies face
The core contribution - using AI to analyze AI usage patterns while preserving privacy - is genuinely novel and represents fresh technical thinking on the privacy-analysis tradeoff. The empirical findings about cross-cultural usage patterns and the methodological approach (defense-in-depth privacy, bottom-up clustering) are original. However, the organizational and governance framing relies heavily on well-established principles (procedural justice, defense-in-depth, cross-functional teams) that are not novel to this field.
Similar to how Google Trends provides aggregate insights about web search behavior without exposing individual queries, CLIO reveals patterns about how AI assistants are used in the real world
The system transforms raw conversations through a multi stage pipeline, extracting key facets like conversation topic or language, clustering semantically similar conversations, generating privacy preserving cluster descriptions
This is a solo academic presentation of a research paper, not a podcast interview with guests. The speaker (presumably Jonathan H. Westover PhD) is presenting their own work rather than being interviewed by a host or discussing with peers. There is no guest caliber to evaluate in the traditional sense - this is a research talk, not a B2B podcast conversation.
Abstract this paper presents Clio CLAUDE Insights and Observations, a privacy preserving platform
Speaker A:
The transcript is exceptionally specific with concrete metrics and named examples throughout: 94% accuracy in ground truth reconstruction, 1 million Claude conversations analyzed, 10%+ web/mobile development, 6-10% research usage, 15-25% coding across platforms, specific companies (Anthropic, OpenAI, Google, Duolingo, Microsoft), real policy examples (election monitoring, prohibited uses), and cross-language data. Minimal hand-waving; nearly all claims include supporting numbers or specific instances.
94% accuracy in reconstructing ground truth topic distributions and achieving undetectable levels of private information in final outputs
coding related tasks dominate usage, with web and mobile application development representing over 10%
This is a solo presentation of academic research with no host-guest dynamic, questions, follow-ups, or productive disagreement. The speaker monologues for 43 minutes without interaction, audience questions, or pushback. While the content is well-organized and moves logically, there is zero conversational engagement, Socratic questioning, or debate that would characterize strong B2B podcast craft. This format fundamentally lacks the conversational elements being evaluated.
Abstract this paper presents Clio CLAUDE Insights and Observations, a privacy preserving platform
Speaker A:
Computed from the transcript - who did the talking, and the words that came up most.
Abstract: This paper presents Clio (Claude insights and observations), a privacy-preserving platform that uses AI assistants to analyze and surface aggregated usage patterns across millions of conversations without requiring human reviewers to read raw user data. The system addresses a critical gap in understanding how AI assistants are used in practice while maintaining robust privacy protections through multiple layers of safeguards. We validate Clio's accuracy through extensive evaluations, demonstrating 94% accuracy in reconstructing ground-truth topic distributions and achieving undetectable levels of private information in final outputs through empirical privacy auditing. Applied to one million Claude.ai conversations, Clio reveals that coding, writing, and research tasks dominate usage, with significant cross-language variations - for example, Japanese conversations discuss elder care at higher rates than other languages. We demonstrate Clio's utility for safety purposes by identifying coordinated abuse attempts, monitoring for unknown risks during high-stakes periods like capability launches and elections, and improving existing safety classifiers.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Abstract this paper presents Clio CLAUDE Insights and Observations, a privacy preserving platform that uses AI assistants to analyze and surface aggregated usage patterns across millions of conversations without requiring human reviewers to read raw user data. The system addresses a critical gap in understanding how AI assistants are used in practice while maintaining robust privacy protections through multiple layers of safeguards. We validate Clio's accuracy through extensive evaluations, demonstrating 94% accuracy in reconstructing ground truth topic distributions and achieving undetectable levels of private information in final outputs through empirical privacy auditing. Applied to 1 million Claude AI uh conversations, Clear reveals that coding, writing, and UM research tasks dominate usage with significant cross language variations. For example, Japanese conversations discuss elder care at higher rates than other languages. We demonstrate clio's utility for safety purposes by identifying coordinated abuse attempts, monitoring for unknown risks during high stakes periods like capability launches and elections, and improving existing safety classifiers. By enabling scalable analysis of real world AI usage while preserving privacy, CLIO provides an empirical foundation for AI safety and governance. Despite widespread interest in ari's societal impact, remarkably little public data exists about how AI assistants are actually used in practice. Which capabilities see real adoption? How does usage vary across cultures, which anticipated benefits and risks materialize in concrete data? This knowledge gap persists despite model providers having access to usage data that could answer these questions, primarily due to four fundamental challenges. First, users share sensitive personal and business information with AI systems, creating tension between privacy protection and the provider's need to understand system usage. Second, having humans review conversations raises ethical concerns due to the repetitive nature of the task and potentially distressing content reviewers might encounter. Third, competitive pressures discourage providers from releasing usage data that could reveal information about their user base to competitors, even when such disclosure would serve the public interest. Finally, the sheer scale millions of daily messages makes manual review impractical. Clio addresses these challenges by using AI assistants themselves to surface aggregated insights across millions m of model interactions while preserving user privacy. Similar to how Google Trends provides aggregate insights about web search behavior without exposing individual queries, CLIO reveals patterns about how AI assistants are used in the real world. The system transforms raw conversations through a multi stage pipeline, extracting key facets like conversation topic or language, clustering semantically similar conversations, generating privacy preserving cluster descriptions, and organizing clusters into a navigable hierarchy. This approach enables discovering both specific patterns of interest and unknown unknowns through an interactive visualization interface. Why Clio Matters now as UH AI systems become more capable and integrated into society, the need for empirical understanding of their real world use intensifies. Pre deployment testing, including red teaming and benchmark evaluations remains crucial but cannot capture all real world usage patterns and emergent risks. Post deployment monitoring provides an essential complement by surfacing patterns that predetermined scenarios might miss. These insights can inform future pre deployment tests, creating a virtuous cycle between empirical observation and proactive safeguards. The timing is particularly critical. AI assistants now handle increasingly sensitive tasks, from medical advice to financial planning to legal research. Without systematic understanding of actual usage, we risk developing governance frameworks disconnected from reality, safety measures that address hypothetical rather than actual risks, um and product improvements that miss what users genuinely need. Clear represents one approach to privacy preserving insight at scale, contributing to an emerging culture of empirical transparency in AI development. The AI Usage Understanding Landscape Defining Privacy Preserving analytics in the AI Context Traditional approaches to understanding technology usage face unique challenges when applied to AI assistants. Unlike analyzing web search queries or application usage logs, AI conversations often contain extended context rich exchanges that may include deeply personal information, proprietary business data, creative works, and um sensitive decision making processes. This richness makes conversations valuable for understanding impact but also heightens privacy concerns. Privacy preserving analytics M must therefore balance two competing objectives generating actionable insights about system usage while protecting individual privacy. Formal privacy frameworks like differential privacy and K anonymity provide strong theoretical guarantees but prove difficult to apply to rich textual outputs. Differential privacy adds noise to query results to prevent inference about individual records, but determining appropriate noise levels for natural language descriptions remains an open challenge. K anonymity ensures each record is indistinguishable from at least K1 others, but defining for complex conversational data lacks clear metrics. CLIO takes a defense in depth approach, implementing multiple privacy layers that collectively reduce private information exposure to undetectable levels in empirical evaluations. This statistical validation approach complements formal guarantees by measuring actual privacy preservation in practice rather than relying solely on theoretical bounds. The system explicitly defines private information broadly encompassing not just individual identifiers but also information that could identify small groups or specific organizations, recognizing that group privacy violations can be as concerning as individual ones. State of Current Approaches to AI Usage Analysis Existing approaches to understanding AI usage fall into several categories, each with limitations that motivate clio's development. Public datasets like wildchat and LMSYS UM Chat provide valuable windows into AI usage but suffer from selection bias. They capture users willing to interact with AI through specific platforms, often for free access rather than representative samples of mainstream usage. These datasets reveal important patterns such as coding dominating 15 to 25% of conversations across platforms, but cannot answer questions about how paid users behave differently or how usage evolves over time on production systems. Academic research on crowd worker generated datasets like the Anthropic Red Team Dataset and Stanford Human Preferences Dataset offers controlled insights into specific scenarios but lacks the ecological validity of real world usage. Crowdworkers following instructions to explore model capabilities produce different interaction patterns than users genuinely trying to accomplish tasks. This gap between research scenarios and actual usage limits the applicability of findings to real world safety and product decisions. Some providers share high level usage statistics or case studies, but these typically offer limited granularity and lack systematic methodology for identifying patterns. Anecdotal evidence from user forums and social media provides qualitative insights, but cannot quantify prevalence or identify unknown patterns. At scale, the field has lacked a systematic, privacy preserving approach to analyzing production usage data, the gap CLIO aims to fill Privacy, Ethics, and Competitive Dynamics the reluctance to analyze and share usage data reflects genuine tensions rather than mere unwillingness. Privacy concerns are paramount. Users reasonably expect their uh, conversations to remain confidential, and analyzing them even in aggregate requires careful consideration. Ethical issues extend beyond privacy to worker well being. Human reviewers who manually examine potentially disturbing content can experience psychological harm, um, raising questions about the ethics of such review processes. Competitive dynamics create additional complexity. Usage patterns reveal valuable information about which features drive engagement, which user segments find value, and which competitors might be gaining traction. Sharing this intelligence could advantage competitors while potentially harming the company's ability to invest in AI safety research through reduced competitive position. However, this competitive concern must be weighed against the public interest in understanding ari's societal impact. CLIO attempts to navigate these tensions by prioritizing privacy through technical safeguards, reducing human exposure to potentially disturbing content through AI mediated analysis, and sharing insights that serve the public interest even when they might reveal competitively sensitive information. This approach recognizes that model providers have both capabilities and responsibilities that extend beyond immediate commercial interests. Organizational and Individual Consequences of AI Assistant Usage Understanding Organizational Adoption Patterns Organizations adopt AI assistance across a remarkably diverse range of use cases, from automating routine tasks to augmenting complex decision making. Clio's analysis reveals that coding related tasks dominate usage, with web and mobile application development representing over 10% of all conversations in our Claude Eye sample. This finding suggests that AI assistants have achieved significant penetration in software development workflows, potentially accelerating development cycles and lowering barriers to entry for programming. Writing and communication tasks comprise another major category, including professional email drafting, document creation, and content editing. This usage pattern indicates that AI assistants serve as cognitive tools for knowledge workers, potentially increasing productivity in communication, intensive roles, research, and educational uses. Representing 6 to 10% of usage suggest that AI assistants function as learning aids and research accelerators, raising important questions about their impact on educational outcomes and research quality. The prevalence of specific use cases varies across linguistic communities, revealing cultural and contextual factors that shape AI adoption. Japanese and Chinese conversations show elevated rates of elder care discussions compared to other languages, potentially reflecting demographic challenges these societies face. Spanish conversations show higher prevalence of economic theory discussions, while anime and UH manga content creation features prominently in Japanese conversations. These cross cultural patterns suggest that AI assistant adoption reflects and potentially amplifies existing social priorities and cultural practices. Individual and Stakeholder Impacts beyond organizational efficiency, AI assistant usage affects individual users in complex ways. The diversity of personal use cases, from dream interpretation to Dungeons and Dragons game mastering to hairstyle advice, demonstrates that AI assistants increasingly mediate intimate and creative aspects of human life. This mediation raises important questions about autonomy, authenticity, and the evolution of human capabilities. Educational uses present particularly significant implications. When students use AI assistance for homework help or exam preparation, the line between legitimate learning support and academic dishonesty becomes ambiguous. Clio's ability to identify clusters of conversations about academic cheating and avoiding detection highlights this tension, suggesting that meaningful numbers of users attempt to use AI assistance in ways that undermine educational integrity. The prevalence of such behavior and how it varies across educational levels and subjects remains an important area for ongoing monitoring and research. Creative and professional tasks increasingly involve AI collaboration, raising questions about attribution, skill development, and the nature of expertise. When users rely on AI assistance for code debugging, legal research, or medical information synthesis, they gain access to capabilities that would otherwise require extensive training or expensive professional services. This democratization of expertise brings both benefits, reduced barriers to accomplishment and risks, potential erosion of professional standards, and increased likelihood of errors. When users lack domain knowledge to evaluate AI outputs critically, the emotional and psychological dimensions of AI usage warrant attention as well. Conversations about dreams, consciousness, and philosophical questions suggest that some users engage AI assistants as interlocutors for existential reflection. While potentially valuable for self exploration, such usage raises questions about the appropriate boundaries of AI involvement in human meaning making and whether AI assistance might substitute for human connection in concerning ways. Evidence based organizational responses, transparent communication about AI capabilities, and limitations. Organizations deploying AI assistance must communicate clearly about system capabilities, limitations, and appropriate use cases. Anthropic's usage policy explicitly prohibits certain uses, including political campaigning, election interference, and generating sexually explicit content, providing clear boundaries for acceptable usage. However, policy alone uh, proves insufficient. Organizations must also educate users about why certain uses pose risks and how to evaluate whether their intended use case falls within acceptable bounds. Effective communication strategies, contextual guidance providing just in time Explanations when users attempt potentially problematic tasks Explaining why certain requests might be refused or flagged. Capability Transparency clearly documenting what the system can and cannot reliably do, including known failure modes and areas of uncertainty. Privacy Education Helping users understand what data is collected, how it's used, and what privacy protections exist. Evolving limitations Regularly updating users as model capabilities change Ensuring they don't rely on outdated mental models of system behavior Anthropic's approach to election monitoring exemplifies transparent communication in action. During the 2024 U.S. general elections, the company used CLIO to identify election related conversations and flag clusters that might indic policy violations rather than silently removing violating content. Anthropic explains to users why certain election related uses, like generating campaign materials, violate policy while others, like learning about voting procedures, are acceptable. This transparency helps users develop accurate mental models of appropriate use. Procedural justice in Content Moderation When AI systems identify potentially violating behavior, organizations must respond in ways that users perceive as, uh, fair and, um, legitimate. Procedural justice the fairness of the processes used to make decisions proves as important as the substantive outcomes themselves. Users who understand why their account was restricted and have opportunities to appeal are more likely to accept enforcement actions as legitimate. Key elements of procedural justice Explanation and transparency Providing clear reasons for enforcement actions rather than opaque policy violations Appeal mechanisms Enabling users to contest decisions they believe were made in error Consistency Applying rules uniformly across similar cases to avoid perceptions of arbitrary enforcement Human oversight Ensuring that high stakes decisions like account termination involve human review rather than fully automated processes OpenAI's approach to research access demonstrates procedural justice principles. When researchers request access to usage data for safety research, OpenAI maintains documented criteria for approval, explains decisions, and provides appeal paths. This procedural clarity helps researchers understand requirements and builds trust even when requests are denied. Clio's design reflects procedural justice considerations by not automating enforcement based solely on cluster membership. Instead, clusters flagged uh as concerning trigger manual review by authorized trust and safety team members who examine individual conversations and make contextualized judgments. This human in the loop approach reduces false positive rates while maintaining legitimacy. Capability Building for Responsible AI Usage Organizations benefit from investing in user education about responsible AI usage rather than relying solely on technical controls. Users who understand AI capabilities, limitations, and risks can make better decisions about when and how to use these systems. Effective capability building approaches Domain specific guidance Providing tailored advice for specific use cases Medical advice Legal research Financial planning about appropriate AI involvement and uh, necessary verification steps Critical evaluation training Teaching users to assess AI outputs critically, recognize hallucinations and errors, and know when to seek expert verification ethical frameworks helping users think through questions of attribution, privacy, and appropriate delegation of judgment to AI systems. Best Practice Sharing Facilitating communities of practice where users share effective patterns for AI collaboration Microsoft's approach in enterprise deployments exemplifies capability building. When deploying CoPilot for Microsoft 365, the company provides extensive training materials helping employees understand when AI assistance adds value versus when it might introduce risks. Domain specific guidance addresses common pitfalls like relying on AI for final legal language without attorney review while celebrating productive use patterns. Duolingo's integration of AI tutoring demonstrates capability building in educational contexts. The language learning platform uses AI assistance to provide personalized practice while clearly communicating to learners that AI interactions complement rather than substitute for comprehensive language instruction. This framing helps users develop realistic expectations about what AI tutoring can accomplish. Operating Model and Technical Controls Technical architecture and operating procedures play crucial roles in enabling responsible AI usage. Organizations must design systems that make safe behaviors easy and unsafe behaviors difficult while preserving flexibility for legitimate edge cases. Effective technical controls Rate Limiting Preventing automated abuse by limiting requests per user or account Input Filtering, Blocking, or flagging clearly prohibited content before it reaches the model. Output Filtering Preventing models from generating prohibited content even when users attempt to elicit it. Behavioral monitoring Tracking patterns across multiple conversations to identify coordinated abuse that individual conversations might not reveal. CLIO itself represents an operating model innovation, using AI assistance to identify patterns of violative behavior that would be invisible at the individual conversation level. When CLIO identifies clusters suggesting coordinated abuse like accounts systematically generating SEO spam across many conversations, it enables enforcement against sophisticated attacks that evade simpler detection methods. Anthropic's multilayered safety approach combines technical controls at ah multiple levels. Models receive training and instructions to refuse harmful requests. Classifiers detect and flag problematic conversations even when models initially respond. Rate limits prevent automated abuse. Usage policies provide clear boundaries. Trust and safety teams review flagged content under strict privacy controls. This defense in depth approach recognizes that no single control provides perfect protection. Google's approach to commercial AI deployment demonstrates the importance of technical controls at scale when offering AI capabilities. Through Google Cloud, the company implements quotas, abuse detection systems, and access controls that prevent individual customers from monopolizing resources or using systems for prohibited purposes. These controls balance openness with responsibility, enabling innovation while preventing misuse. Financial and benefit supports for positive use cases Organizations can actively promote beneficial AI usage through strategic pricing, access policies and partnership programs that make AI assistance available for high social value applications. Strategies for supporting positive use cases Educational access Providing free or subsidized access to students, educators, and educational institutions Research Partnerships Enabling academic researchers to access usage data or computational resources for safety and social impact research Non profit support Offering preferential pricing or capabilities to organizations working on social challenges Safety research funding Investing in external research on AI safety, fairness, and beneficial applications Anthropic's approach to research access exemplifies this support model. The company provides researchers studying AI safety with access to models and usage data UH under strict privacy controls, enabling independent analysis that informs both anthropics development and broader community understanding. This investment in external scrutiny demonstrates commitment to safety beyond immediate commercial interests. OpenAI's chatgpt.edu program shows how targeted access can support beneficial use by offering educational institutions specialized access. Designed for learning applications, OpenAI enables exploration of ARI's educational potential while building safeguards against academic dishonesty. The program includes features like activity dashboards that help educators understand how students use AI. Google's AI for Social Good initiatives demonstrate financial support for positive applications by providing computational resources, expertise, and UM funding to organizations addressing social and environmental challenges. Google enables AI UH application to high value problems that might otherwise lack resources for sophisticated AI deployment Building long term capabilities for responsible AI development Empirical observation and Continuous learning Systems the rapid evolution of AI capabilities means that static governance frameworks quickly become outdated. Organizations must build capabilities for continuous empirical observation, learning, and adaptation to maintain effective governance as models and usage patterns evolve. Pleo exemplifies this approach by enabling ongoing analysis of real world usage without requiring predetermined hypotheses about what patterns might emerge. The system's bottom up design Clustering conversations based on semantic similarity rather than predefined categories allows discovery of usage patterns that developers and safety teams might not anticipate. This capability proves especially valuable during periods of uncertainty. New capability launches major world events or rapid changes in model behavior. Building sustainable empirical observation requires scalable analysis. Infrastructure systems that can process millions of conversations efficiently enough to provide timely insights Multilingual capabilities Analysis methods that work across languages to understand global usage patterns Temporal tracking Monitoring how usage patterns evolve over time to identify emerging trends before they become widespread Cross domain integration Combining usage analysis with other data sources like user surveys, external research, and safety incident reports to develop holistic understanding, Organizations should view empirical observation not as one time audit but as continuous monitoring analogous to how technology companies monitor system performance and reliability. Just as engineering teams use observability platforms to detect performance degradation or service failures, safety and governance teams need observability into usage patterns and risks. This requires investment in infrastructure, dedicated teams with appropriate expertise, and organizational processes that translate insights into action the feedback loop between observation and action proves crucial when CLIO identifies concerning patterns like coordinated abuse attempts or classifier false positives. Those insights should trigger concrete responses. Updating safety classifiers, refining usage policies, improving user education, or adjusting model training. Without this action oriented approach, observation generates data without impact. Data Stewardship and Privacy Infrastructure Long term responsible AI UH development requires robust data stewardship policies, processes, and technologies that ensure data is collected, stored, analyzed, and shared in ways that respect privacy while enabling necessary analysis. Clio's privacy architecture demonstrates key principles of effective data stewardship. Multiple privacy layers work in concert. Conversation summaries extract key information while excluding private details. Cluster aggregation thresholds ensure clusters represent many users rather than individuals. Cluster summaries are generated with explicit privacy instructions. An automated auditing removes clusters containing private information. This defense in depth approach recognizes that no single protection provides perfect privacy. Essential elements of privacy infrastructure Privacy by design Building privacy protections into systems from inception rather than adding them as afterthoughts Access controls Limiting who can view sensitive data to authorized personnel with legitimate business needs Audit capabilities Maintaining logs of data access and analysis to ensure accountability Retention policies Deleting data when it no longer serves necessary purposes rather than retaining indefinitely user UH clearly communicating to users what data is collected and how its used. Organizations must also grapple with tensions between privacy and other values. Identifying coordinated abuse requires linking behavior across accounts potentially conflicting with strong anonymization. Improving model safety through analyzing failure cases requires examining specific problematic conversations, creating tension with policies against human review. Navigating these tensions requires explicit value judgments about acceptable tradeoffs rather than pretending conflicts don't exist. The role of differential privacy and formal guarantees in AI usage analysis remains an important area for research and development. While CLIO currently relies on empirical privacy validation rather than formal guarantees, future systems might incorporate differential privacy techniques that provide mathematical bounds on privacy loss. However, the richness of natural language outputs makes direct application challenging. Research into privacy. Preserving text generation could enable stronger formal guarantees while maintaining analytical utility. Distributed Leadership and Cross Functional Collaboration Effective AI governance requires expertise spanning multiple domains machine learning, software, engineering policy and legal analysis, ethics, social science, and domain specific knowledge for particular applications. No single team or individual possesses all necessary expertise, making cross functional collaboration essential. Clio's development reflects distributed leadership. AI researchers developed core clustering algorithms, engineers built scalable infrastructure, safety specialists designed privacy protections, policy experts defined acceptable use cases, social scientists contributed analysis frameworks, and ethics experts helped navigate tensions between competing values. This collaboration produced a system more robust than any single discipline could have created. Organizational structures that enable distributed leadership Cross functional teams bringing together diverse expertise for major initiatives rather than siloing work by discipline Embedded specialists Placing experts in ethics, safety or policy directly within product and research team structured consultation processes Creating mechanisms for seeking input from relevant stakeholders before major decisions Transparent decision making Documenting key decisions and rationales to enable learning and accountability Organizations should resist the temptation to centralize all AI governance decisions in a single team as this creates bottlenecks and fails to leverage domain expertise. Instead, governance should be distributed across the organization with clear accountability for different types of decisions. Product teams might make day to day choices about feature design within guardrails established by safety teams who in turn operate within boundaries set by executive leadership and informed by external feedback. Purpose, Values and Organizational Culture Technical systems and processes ultimately rest on organizational culture, shared values, assumptions and priorities that shape how people make decisions when formal rules don't provide clear guidance. Building culture that prioritizes responsible AI development requires explicit attention to values and mechanisms that reinforce them. Key cultural elements Safety as core value Treating safety as fundamental rather than constraint on innovation Transparency as norm Defaulting to sharing information about capabilities, limitations and usage patterns and less specific harm would result Epistemic humility Acknowledging uncertainty about long term impacts and potential risks we haven't imagined external accountability Welcoming scrutiny from researchers, civil society and the public rather than defending against it. Anthropic's decision to publish this paper exemplifies these values in practice, sharing detailed information about Clio, including its capabilities, limitations and actual usage insights serves the public interest even though it reveals competitively sensitive information about user behavior and potentially advantages competitors. This transparency reflects organizational commitment to empirical grounding of AI governance rather than relying solely on internal judgment. Building and maintaining purpose driven culture requires ongoing investment beyond one time statements of values Hire for values alignment Selecting team members who demonstrate genuine commitment to uh, responsible development Reward values aligned behavior Ensuring that promotion and recognition systems reward long term safety thinking rather than only short term metrics Create space for dissent Enabling team members to raise concerns without fear of retaliation. Learned from failures Treating safety incidents as learning opportunities rather than occasions for blame Engage with critics Seeking out and taking seriously feedback from those skeptical of AI development the field of AI development faces profound challenges in aligning rapid capability growth with robust safety and governance. No single organization will solve these challenges alone. Progress requires collaboration, knowledge sharing and um, willingness to prioritize collective benefit over individual competitive advantage. By sharing CLIO and its insights openly, we hope to contribute to this collaborative effort. Conclusion CLIO demonstrates that privacy preserving analysis of real world AI usage is both technically feasible and practically valuable through multilayered privacy protections. The system surfaces meaningful insights from coding dominating usage patterns to cross cultural variations in application focus to coordinated abuse attempts while maintaining user uh privacy at empirically validated levels. The platform enables discovery of unknown unknowns that predetermined tests might miss, complementing proactive safety measures with reactive learning from actual deployment. The findings shared in this top use cases Multilingual patterns safety classifier performance represent early explorations of questions that will only grow more important as AI systems become more capable and widespread. Understanding that Japanese users discuss elder care at elevated rates, that coding tasks dominate across languages and platforms, or that certain clusters of conversations systematically attempt to evade safety measures provides concrete grounding for governance decisions that might otherwise rely on speculation. Looking ahead, several research directions warrant attention. Extending clio's analysis to long term usage trajectories could reveal how individuals relationships with AI assistants evolve over time. Developing more sophisticated multilingual analysis could uncover cross cultural patterns invisible in English only studies. Studies Improving formal privacy guarantees while maintaining analytical utility remains an important technical challenge, and expanding beyond conversational data to include outcomes. How AI assisted work products differ from unassisted ones would provide crucial insight into real world impacts. We share CLIO not as a complete solution but as one approach to an urgent challenge, grounding AI safety and governance in empirical reality rather than hypothetical speculation. The platform's effectiveness depends on continued refinement based on new insights, evolving capabilities, and feedback from the broader research community. We invite that engagement, recognizing that responsible AI development requires sustained collaboration across organizations, disciplines, and perspectives. By making AI usage patterns visible while preserving privacy, systems like CLIO can help ensure that governance frameworks, safety measures, and product improvements respond to actual usage rather than imagined scenarios. This empirical grounding, combined with proactive safety.