# Overview Source: https://documentation.5x.co/core-features Explore the powerful capabilities that make 5X your complete data platform solution **Focus:** Understand the full capabilities of the 5X platform and how each feature contributes to your data strategy. The 5X platform provides a comprehensive suite of integrated features designed to handle every aspect of your data workflow - from ingestion and warehousing to AI-powered insights and conversational interfaces. **Integrated by Design** All core features work seamlessly together, sharing data models, security policies, and orchestration workflows. This integration eliminates the complexity of managing multiple tools and vendors. ## Core platform features Explore each feature to understand how it fits into your data architecture: **Universal data connectivity** Connect to 600+ data sources with real-time and batch ingestion capabilities. **Integrated development environment** Build data pipelines, models, and applications with dbt, Python, and Streamlit. **Workflow automation** Schedule and manage complex data workflows with dependency tracking and monitoring. **Interactive dashboards & reports** Create compelling visualizations and self-service analytics for your stakeholders. **Metrics data models** Build and manage Cube.js metrics layer projects with Git-based version control and BI integration. **Custom applications** Build and deploy data-driven applications with embedded analytics and AI. **Natural language interface** Query your data and get insights using plain English conversations. ## How features work together The 5X platform is designed as an integrated ecosystem where each feature enhances the others: ### **Data foundation** * **Ingestion** continuously feeds fresh data into your warehouse * **IDE** provides the development environment to build models, applications, and data pipelines ### **Automation & insights** * **Orchestration** ensures your data pipelines run reliably and efficiently * **Metrics Store** creates consistent definitions across all downstream use cases and provides Cube.js metrics layer with project management and BI integration * **Business Intelligence** delivers insights to business stakeholders ### **Advanced applications** * **Data & AI Apps** embed analytics directly into business workflows * **Conversational AI** makes data accessible to everyone through natural language **Start with quickstart** Get your workspace running and see core features in action in 15 minutes. # App connections Source: https://documentation.5x.co/core-features/business-intelligence/app-connections Configure Business Intelligence app connections to securely connect to your data warehouse and metrics layer **Focus:** Learn how to set up and manage Business Intelligence app connections that enable 5X BI to access your data warehouse and metrics layer securely. ## App connections overview App Connections are 5X's integrated approach to data source management. Unlike traditional BI tools where database connections are managed within the BI interface, 5X Business Intelligence uses App Connections for centralized, secure data access. ### **Why App Connections matter** **Unified control** Manage all data connections from one place across your entire 5X platform. **Consistent security** Apply consistent security policies and access controls across all platform capabilities. **Easy configuration** Pre-configured connection templates make setup faster and more reliable. **Seamless integration** App Connections work seamlessly with ingestion, IDE, orchestration, and other 5X capabilities. ## Understanding BI app connections ### **What are Business Intelligence app connections?** Business Intelligence app connections define how 5X BI accesses your data warehouse. They contain all the technical details needed for BI to query your data, including: * **Authentication credentials** - Username, password, or key pair authentication * **Connection parameters** - Database, schema, warehouse, and role settings * **Security settings** - Access permissions and row-level security policies * **Performance configuration** - Timeout settings, connection pooling, and optimization ### **Two connection options** 5X Business Intelligence supports two types of data connections: **Direct warehouse access:** * Connect directly to your Snowflake, BigQuery, or other warehouse * Full SQL capabilities with complete access to all data and SQL functions * Real-time queries that access data as it exists in your warehouse * Custom transformations using custom SQL for complex analysis **Direct Metrics Store project access:** * Connect 5X BI directly to your 5X Metrics Store projects * Unified metrics with access to pre-defined business metrics and KPIs * Consistent definitions using standardized metric definitions across all tools * Simplified queries for business metrics without complex SQL * Governance through centralized metric management and approval processes 5X Business Intelligence App Connection Configuration ## Creating BI app connections ### **Step 1: Access app connections** 1. **Navigate to Settings** * From your workspace, click **"Settings"** in the left sidebar * Select **"App connections"** from the settings menu 2. **Create new connection** * Click **"+ New connection"** * Select **"Business Intelligence"** as the connection type ### **Step 2: Configure connection details** Set connection name, description, and tags for easy identification and management. Choose between "Warehouse" for direct access or "Metrics Store" for unified metrics. Configure authentication method (username/password or key pair) with appropriate credentials. Set database, schema, warehouse, and role settings based on your data architecture. ### **Step 3: Test and save connection** 1. **Test connection** - Verify credentials and access permissions 2. **Review settings** - Confirm all configuration details are correct 3. **Save connection** - Create the app connection for use in BI ## Connection management ### **Troubleshooting common issues** **Possible causes:** * Incorrect username or password * Expired or invalid credentials * Wrong authentication method * Account locked or disabled **Solutions:** * Verify credentials with database administrator * Check account status and permissions * Confirm authentication method configuration * Reset credentials if necessary **Possible causes:** * Insufficient database permissions * Missing role assignments * Schema or table access restrictions * Row-level security policies **Solutions:** * Review role permissions and assignments * Check schema and table access rights * Verify row-level security configuration * Contact database administrator for access **Possible causes:** * Network connectivity issues * Database server overload * Firewall or proxy restrictions * Incorrect connection parameters **Solutions:** * Test network connectivity * Check database server status * Verify firewall and proxy settings * Review connection timeout configurations ### **Updating connections** **When to update:** * **Credential rotation** - Regular security credential updates * **Permission changes** - Modified access requirements * **Infrastructure changes** - Database or network modifications * **Performance optimization** - Improved connection settings **Update process:** 1. **Access connection settings** - Navigate to app connections 2. **Edit connection** - Modify necessary settings 3. **Test changes** - Verify updated configuration 4. **Save updates** - Apply changes to the connection ## Security best practices ### **Access control** **Minimal permissions** Grant only the minimum permissions necessary for BI functionality. **Dedicated roles** Use dedicated database roles specifically for BI access. **Access reviews** Regularly review and audit access permissions and usage patterns. **Credential management** Use secure credential storage and rotation practices. ## Next steps With your Business Intelligence app connection configured, you're ready to start building dashboards and exploring your data: **Next: Create dashboards** Once your app connection is set up, learn how to build interactive dashboards. **Explore your data** Learn how to explore datasets and write custom SQL queries. # Chart building Source: https://documentation.5x.co/core-features/business-intelligence/chart-building Master the art of creating compelling charts and visualizations from your data with 50+ chart types **Focus:** Learn how to build effective charts and visualizations that tell compelling data stories and enable users to discover insights. ## Chart building overview Charts are the foundation of effective data visualization in 5X Business Intelligence. With over 50 chart types available, you can create visualizations that match your data and audience needs, from simple bar charts to complex geospatial maps. ### **Choosing the right chart type** The key to effective data visualization is selecting the appropriate chart type for your data and message: **Trends and patterns over time** Use line charts, area charts, or time-series bar charts to show how metrics change over time. **Comparing groups or categories** Bar charts, column charts, and pie charts work well for comparing different categories or groups. **Location-based insights** Maps and geospatial charts help visualize data tied to specific locations or regions. **Correlations and connections** Scatter plots, bubble charts, and network diagrams reveal relationships between variables. ## Creating your first chart ### **Step 1: Access the chart builder** 1. **Navigate to Business Intelligence** * From your workspace, click **"BI"** in the left sidebar and navigate to the **"Charts"** tab. * Click **"+ Chart"** in the top right corner 2. **Choose your data source** * Select from available datasets * Choose between warehouse data or metrics layer metrics * Configure any necessary data transformations 5X Business Intelligence Chart Builder Interface ### **Step 2: Configure chart data** Choose from 50+ chart types based on your data and visualization goals. Preview how your data will look with each type. Set up categorical data (groups, categories, time periods) that will structure your visualization. Define the numerical values you want to visualize (counts, sums, averages, percentages). Add filters to focus on specific data subsets or time periods relevant to your analysis. 5X Business Intelligence Chart Builder Interface ## Chart types and use cases **Reference guide:** Explore 50+ chart types organized by category, with detailed use cases and configuration guidance for each visualization type. * **Use cases:** Comparing values across categories * **Best for:** Sales by region, product performance, survey results * **Configuration:** Category on one axis, values on the other * **Use cases:** Showing trends over time * **Best for:** Revenue trends, user growth, performance metrics * **Configuration:** Time on x-axis, metrics on y-axis * **Use cases:** Showing parts of a whole * **Best for:** Market share, budget allocation, survey responses * **Configuration:** Categories as slices, values as percentages * **Use cases:** Showing cumulative values over time * **Best for:** Stacked metrics, cumulative growth, layered data * **Configuration:** Time series with filled areas * **Use cases:** Showing relationships between two variables * **Best for:** Correlation analysis, outlier detection, clustering * **Configuration:** Two numerical axes with optional size/color encoding * **Use cases:** Showing relationships with additional dimensions * **Best for:** Market analysis, performance comparisons, multi-dimensional data * **Configuration:** X/Y axes plus bubble size and color * **Use cases:** Showing patterns in large datasets * **Best for:** User behavior analysis, performance matrices, correlation tables * **Configuration:** Two categorical axes with color intensity encoding * **Use cases:** Showing hierarchical data with size relationships * **Best for:** Budget allocation, organizational structure, market analysis * **Configuration:** Hierarchical categories with size encoding * **Use cases:** Global data visualization * **Best for:** International sales, user distribution, market penetration * **Configuration:** Geographic coordinates with color/size encoding * **Use cases:** Regional data analysis * **Best for:** Regional performance, demographic analysis, territory management * **Configuration:** Administrative boundaries with data encoding * **Use cases:** Specific geographic areas or custom territories * **Best for:** Store locations, service areas, custom regions * **Configuration:** Custom geographic data with visualization encoding ## Data configuration ### **Dimensions and metrics** Understanding dimensions and metrics is fundamental to creating effective charts. These two data types work together to structure your visualizations and provide meaningful insights. **Dimensions (categorical data):** Dimensions are the categorical variables that define how your data is grouped and organized. They provide the structure and context for your visualizations. **Temporal data that shows trends and patterns over time** Time dimensions are essential for analyzing trends, patterns, and changes over different time periods. They help you understand how your metrics evolve and identify seasonal patterns or growth trends. * **Examples:** Date, month, quarter, year, hour, day of week * **Use cases:** Revenue trends, user activity patterns, seasonal analysis * **Best practices:** Choose appropriate granularity (daily vs monthly) based on your data volume and analysis needs **Discrete categories that segment your data** Categorical dimensions help you segment and compare your data across different groups or categories. They're perfect for identifying top performers, comparing regions, or analyzing different product categories. * **Examples:** Product categories, geographic regions, customer segments, departments * **Use cases:** Comparing performance across different groups, identifying top performers * **Best practices:** Limit categories to 5-7 for readability, use "Other" for smaller groups **Multi-level categorical data with parent-child relationships** Hierarchical dimensions allow you to organize data in a structured, drill-down format. They're ideal for geographic analysis, organizational reporting, and any data that has natural parent-child relationships. * **Examples:** Country > State > City, Product > Category > Subcategory * **Use cases:** Drill-down analysis, organizational reporting, geographic analysis * **Best practices:** Design logical drill paths that match your business structure **Calculated categories and groupings created from existing data** Custom dimensions let you create business-specific categorizations that aren't directly available in your source data. They're perfect for creating performance tiers, age groups, or any custom segmentation your business needs. * **Examples:** Age groups, performance tiers, custom segments * **Use cases:** Creating business-specific categorizations, simplifying complex data * **Best practices:** Use clear naming conventions and document calculation logic *** **Metrics (numerical data):** Metrics are the quantitative values you want to measure, analyze, and visualize. They represent the "what" you're measuring in your charts. **Simple counting of records, events, or occurrences** Count metrics are fundamental for measuring volume, activity, and frequency. They help you understand how many times something happens, how many items exist, or how many events occur within a given period. * **Examples:** Number of orders, page views, customer registrations, support tickets * **Use cases:** Volume analysis, activity tracking, performance monitoring * **Best practices:** Use for discrete events, ensure consistent counting logic **Aggregated totals of numerical values** Sum metrics are essential for financial reporting and resource allocation. They help you understand total values, cumulative amounts, and aggregate performance across different categories or time periods. * **Examples:** Total revenue, gross sales, inventory quantities, hours worked * **Use cases:** Financial reporting, resource allocation, performance totals * **Best practices:** Consider currency formatting, handle null values appropriately **Mean values and performance indicators** Average metrics provide insights into typical performance, efficiency, and quality levels. They help you benchmark performance and understand the central tendency of your data. * **Examples:** Average order value, customer satisfaction scores, response times * **Use cases:** Performance benchmarking, quality metrics, efficiency analysis * **Best practices:** Consider outliers, use appropriate decimal precision **Calculated percentages, rates, and proportions** Ratio metrics are powerful for understanding relationships, efficiency, and comparative performance. They help you measure success rates, growth, and relative performance across different segments. * **Examples:** Conversion rates, profit margins, market share, growth rates * **Use cases:** Performance ratios, efficiency metrics, comparative analysis * **Best practices:** Use consistent denominators, format as percentages when appropriate ### **Advanced data configuration** Advanced data configuration allows you to transform, calculate, and manipulate your data to create more meaningful and insightful visualizations. These features enable you to derive new insights from existing data without modifying your underlying data sources. **Calculated fields:** Calculated fields let you create new metrics and dimensions by combining, transforming, or analyzing existing data fields. This is particularly useful when you need metrics that aren't directly available in your source data. **Create new metrics using mathematical operations on existing data** Custom formulas allow you to create sophisticated business metrics by combining existing data fields with mathematical operations. This is essential for creating financial ratios, performance indicators, and business-specific calculations. **Examples:** * Profit margin: `(Revenue - Cost) / Revenue * 100` * Growth rate: `(Current Period - Previous Period) / Previous Period * 100` * Customer lifetime value: `Average Order Value * Purchase Frequency * Customer Lifespan` **Use cases:** Financial ratios, performance indicators, business-specific metrics **Best practices:** Use descriptive names, validate formulas with sample data, document calculation logic **Create categorical fields using IF/THEN statements for data categorization** Conditional logic enables you to create categorical dimensions and metrics based on business rules and thresholds. This is perfect for customer segmentation, performance categorization, and implementing complex business logic. **Examples:** * Customer tier: `IF(Total_Spent > 1000, "Premium", IF(Total_Spent > 500, "Standard", "Basic"))` * Performance rating: `IF(Score >= 90, "Excellent", IF(Score >= 70, "Good", "Needs Improvement"))` * Season classification: `IF(MONTH(Date) IN (12,1,2), "Winter", IF(MONTH(Date) IN (3,4,5), "Spring", ...))` **Use cases:** Customer segmentation, performance categorization, business rule implementation **Best practices:** Keep logic simple and readable, test edge cases, use consistent naming **Perform calculations using basic arithmetic operations** Mathematical operations provide the foundation for creating derived metrics through basic arithmetic. These operations are essential for unit economics, efficiency calculations, and comparative metrics. **Examples:** Add, subtract, multiply, divide values to create derived metrics **Use cases:** Unit economics, efficiency calculations, comparative metrics **Best practices:** Handle division by zero, consider data types, validate results **Perform time-based calculations and comparisons** Date calculations are crucial for analyzing customer lifecycle, calculating trends, and creating time-based metrics. These functions help you understand temporal patterns and relationships in your data. **Examples:** * Days since last purchase: `DATEDIFF(CURRENT_DATE, Last_Purchase_Date)` * Quarter over quarter growth: `(Q2_Revenue - Q1_Revenue) / Q1_Revenue` * Age in years: `DATEDIFF(CURRENT_DATE, Birth_Date) / 365` **Use cases:** Customer lifecycle analysis, trend calculations, time-based metrics **Best practices:** Consider time zones, handle leap years, use appropriate date functions **Data transformations:** Data transformations help you aggregate, group, and manipulate data to better suit your analysis needs. These operations are essential for creating meaningful visualizations from raw data. **Combine multiple values into single summary statistics** Aggregation functions are fundamental for summarizing data and creating meaningful metrics. They help you understand overall performance, trends, and patterns by combining multiple data points into single summary statistics. **Functions:** * **SUM:** Total values across groups (total revenue, total units sold) * **COUNT:** Count occurrences (number of orders, customer count) * **AVG:** Calculate averages (average order value, mean response time) * **MIN/MAX:** Find extreme values (lowest price, highest score) **Use cases:** Summary reporting, performance metrics, trend analysis **Best practices:** Choose appropriate aggregation level, handle null values, consider data distribution **Perform calculations across related rows without grouping** Window functions enable advanced analytical calculations that operate across related rows without collapsing data into groups. They're essential for trend analysis, rankings, and comparative calculations. **Functions:** * **Running totals:** Cumulative sums over time periods * **Moving averages:** Rolling averages for trend smoothing * **Rankings:** Position within groups (top performers, percentile rankings) **Use cases:** Trend analysis, performance rankings, comparative analysis **Best practices:** Define appropriate window frames, consider performance implications **Organize continuous data into discrete categories** Grouping and binning help you organize continuous or large datasets into manageable, discrete categories. This is essential for customer segmentation, performance categorization, and data simplification. **Examples:** * **Age groups:** 18-25, 26-35, 36-45, 46+ * **Revenue tiers:** $0-10K, $10K-50K, $50K-100K, $100K+ * **Performance bands:** Low, Medium, High **Use cases:** Customer segmentation, performance categorization, data simplification **Best practices:** Use meaningful boundaries, ensure adequate sample sizes per group **Focus on specific data subsets and organize results** Filtering and sorting operations help you focus your analysis on relevant data subsets and organize results in meaningful ways. These operations are crucial for targeted analysis and data exploration. **Operations:** * **Date range filtering:** Focus on specific time periods * **Value filtering:** Include only relevant data ranges * **Top N filtering:** Show only top performers or categories **Use cases:** Focused analysis, performance monitoring, data exploration **Best practices:** Apply filters consistently, document filter logic, consider impact on sample size ### **Custom SQL** Custom SQL provides unlimited flexibility for advanced users who need to go beyond the standard chart building interface. This powerful feature allows you to write custom queries that can handle complex data transformations, advanced analytics, and sophisticated business logic. **When to use custom SQL:** Custom SQL is ideal when you need capabilities that go beyond the standard chart builder interface. * **Complex calculations** - Multi-step data transformations that require advanced SQL functions * Examples: Cohort analysis, customer lifetime value calculations, complex financial ratios * Use cases: Advanced analytics, custom business metrics, sophisticated reporting * Benefits: Full control over calculation logic, ability to use advanced SQL functions * **Advanced filtering** - Sophisticated WHERE clauses with complex conditions * Examples: Multi-condition filters, subqueries for filtering, dynamic date ranges * Use cases: Complex data segmentation, conditional analysis, dynamic reporting * Benefits: Precise control over data selection, ability to use subqueries and CTEs * **Joins and unions** - Combining multiple data sources for comprehensive analysis * Examples: Customer data joined with transaction data, multiple product tables combined * Use cases: Cross-system analysis, comprehensive reporting, data integration * Benefits: Access to related data, ability to create unified views * **Performance optimization** - Optimized queries for large datasets and complex operations * Examples: Pre-aggregated data, optimized joins, efficient subqueries * Use cases: Large-scale analytics, performance-critical dashboards, real-time reporting * Benefits: Better query performance, reduced load times, efficient resource usage ## Troubleshooting ### **Common chart issues** **Possible causes:** * Data source connection issues * Incorrect dimension/metric configuration * Data filtering issues * Query errors or timeouts **Solutions:** * Verify data source connections * Check dimension and metric settings * Review applied filters * Test queries independently **Possible causes:** * Color scheme conflicts * Font or sizing issues * Label overlap or truncation * Responsive design problems **Solutions:** * Adjust color schemes and contrast * Modify font sizes and spacing * Optimize label positioning * Test on different screen sizes **Possible causes:** * Large dataset volumes * Inefficient queries * Complex calculations * Network latency issues **Solutions:** * Optimize data queries and filters * Implement data aggregation strategies * Use caching for frequently accessed data * Monitor and optimize infrastructure *** **Next: Explore data** Learn how to explore datasets and write custom SQL queries. **Create dashboards** Combine your charts into compelling interactive dashboards. # Dashboard creation Source: https://documentation.5x.co/core-features/business-intelligence/dashboard-creation Learn how to create compelling interactive dashboards with charts, filters, and collaborative features **Focus:** Master the art of creating interactive dashboards that tell compelling data stories and enable self-service analytics for your organization. ## Dashboard creation overview Dashboards in 5X Business Intelligence are powerful tools for presenting data insights in an organized, interactive format. They combine multiple charts, filters, and text elements to create comprehensive views of your business metrics and KPIs. ### **What makes a great dashboard?** **Focused objectives** Each dashboard should have a clear purpose and target audience. Avoid trying to show everything in one view. **Organized structure** Arrange charts in a logical flow that guides users through your data story from high-level metrics to detailed insights. **User engagement** Use filters, drill-downs, and cross-filtering to let users explore data and find their own insights. **Professional appearance** Maintain consistent colors, fonts, and styling to create a professional, branded experience. ## Creating your first dashboard ### **Step 1: Access the dashboard builder** 1. **Navigate to Business Intelligence** * From your workspace, click **"BI"** in the left sidebar * Click **"+ Dashboard"** in the top right corner 2. **Start building** * Begin with a blank canvas to build your custom dashboard layout 5X Business Intelligence Dashboard Creation Interface ### **Step 2: Configure dashboard settings** Set your dashboard name, description, and tags for easy organization and discovery. Set up automatic data refresh intervals to keep your dashboard current with the latest data. Save your dashboard configuration and proceed to the layout builder. ### **Step 3: Build your dashboard layout** The dashboard builder uses a flexible grid system for arranging charts and components: **Adding components:** * **Charts** - Drag existing charts or create new ones * **Text boxes** - Add titles, descriptions, and context * **Filters** - Create global filters for dashboard-wide filtering * **Images** - Include logos, diagrams, or visual elements **Layout principles:** * **Top-down flow** - Place most important metrics at the top * **Group related content** - Use rows and columns to organize related charts * **White space** - Allow breathing room between elements * **Responsive design** - Ensure dashboards work on different screen sizes ## Dashboard components ### **Charts and visualizations** Charts are the core building blocks of your dashboard: **Chart types available:** * **Time series** - Line charts, area charts for trend analysis * **Categorical** - Bar charts, pie charts for comparisons * **Geographic** - Maps for location-based data * **Tables** - Detailed data views with sorting and filtering * **Custom** - Advanced visualizations for specific use cases **Adding charts to dashboards:** 1. Access the dashboard builder interface 2. Select charts from your existing chart library 3. Position and resize the chart in your layout 4. Configure chart-specific settings and interactions ### **Dashboard filters** Filters enable users to interact with your dashboard data: **Filter types:** * **Date range** - Filter data by time periods * **Categorical** - Filter by specific values or categories * **Numeric** - Filter by value ranges or thresholds * **Custom SQL** - Advanced filtering with custom logic **Filter configuration:** * **Scope** - Apply to specific charts or entire dashboard * **Default values** - Set sensible defaults for new users * **Required filters** - Force users to select certain filters * **Cascading filters** - Create dependent filter relationships ### **Text and annotations** Enhance your dashboard with contextual information: **Text elements:** * **Titles and headers** - Clear section identification * **Descriptions** - Explain metrics and provide context * **Instructions** - Guide users on how to interact with the dashboard * **Annotations** - Highlight important insights or changes **Best practices:** * Keep text concise and actionable * Use consistent formatting and styling * Provide context for complex metrics * Include data refresh timestamps ## Best practices ### **Design principles** **Know your audience** Design dashboards for specific user personas and their information needs. **Layer information** Start with high-level metrics and allow users to drill down for details. **Standardized KPIs** Use consistent metric definitions across all dashboards for reliable insights. **Keep data fresh** Ensure dashboards reflect current data with appropriate refresh schedules. ### **Common pitfalls to avoid** **Overwhelming users:** * Too many charts on one dashboard * Complex layouts without clear hierarchy * Missing context or explanations * Inconsistent styling and formatting **Performance issues:** * Loading too much data at once * Inefficient queries and calculations * Missing data refresh strategies * Ignoring mobile user experience **Poor user experience:** * Unclear navigation and interactions * Missing or confusing filters * Inconsistent metric definitions * Lack of responsive design ## Troubleshooting ### **Common dashboard issues** **Possible causes:** * Data source connection issues * Insufficient permissions for data access * Query errors or timeouts * Missing or invalid chart configurations **Solutions:** * Verify data source connections * Check user permissions and roles * Review query performance and optimization * Validate chart configuration settings **Possible causes:** * Incorrect filter configuration * Data type mismatches * Missing filter dependencies * Cache or refresh issues **Solutions:** * Review filter settings and scope * Verify data types and formats * Check filter dependencies and cascading * Clear cache and refresh data **Possible causes:** * Large data volumes * Inefficient queries * Complex calculations * Network or infrastructure issues **Solutions:** * Optimize data queries and filters * Implement data aggregation strategies * Use caching for frequently accessed data * Monitor and optimize infrastructure *** **Next: Create charts** Learn how to build compelling charts and visualizations for your dashboards. **Connect to data** Set up your data connections to power your dashboards. # Dashboard embedding Source: https://documentation.5x.co/core-features/business-intelligence/dashboard-embedding Embed 5X BI dashboards into external applications with secure authentication and row-level security **Focus:** Learn how to embed 5X Business Intelligence dashboards into your applications with secure authentication, access control, and customizable user experiences. ## Dashboard embedding overview Dashboard embedding enables you to integrate 5X Business Intelligence dashboards directly into your own applications via iframe. This provides a seamless way to bring data analytics into your native environment while maintaining security and access control. ### **Why embed dashboards?** **Seamless integration** Provide analytics within your application's native interface for a cohesive user experience. **Stay in your app** Keep users in your application instead of redirecting them to external BI tools. **Brand consistency** Maintain your application's look and feel while providing powerful analytics. **Easy deployment** Deploy analytics capabilities without building custom visualization components. ## Getting started with embedding ### **Two-step process** Dashboard embedding is organized into two main phases: **Phase 1: Preparation** * Collect required 5X BI asset information * Configure embedding settings and permissions * Gather technical details for integration **Phase 2: Deployment** * Install the Superset Embedded SDK * Implement guest token generation * Embed dashboards in your application ### **Prerequisites** **Required permissions:** * **Admin/Developer role** access in your 5X workspace * **Dashboard access** - Ability to view and configure dashboards * **API access** - 5X API key for guest token generation To obtain your 5X API key, please reach out to 5X Support at [support@5x.co](mailto:support@5x.co), contact us via Intercom, or reach out to your dedicated customer success manager. **Technical requirements:** * **Backend service** - For secure guest token generation * **Frontend application** - For SDK integration and iframe embedding * **HTTPS** - Secure connection for embedded dashboards ## Phase 1: Preparation ### **Enable dashboard embedding** 1. **Access your dashboard** * Navigate to the dashboard you want to embed * Click the options menu (three dots) in the dashboard header * Select **"Embed dashboard"** 2. **Configure allowed domains** * Enter the domains where the dashboard will be embedded * Include development and production domains * Use full URLs with protocol (https\:// or http\://) **Domain configuration tips:** * Wildcards are not supported - specify exact domains * Include all environments (dev, staging, prod) * Leave empty to allow embedding from any domain (not recommended for production) 3. **Enable embedding** * Click **"Enable Embedding"** * Save the provided **Embedded Dashboard ID** for use in your application ### **Collect required information** **Embedded Dashboard ID:** * Unique identifier for the embedded dashboard * Used with the Superset SDK to embed the dashboard * Required for guest token generation **Dashboard configuration:** * **Dashboard settings** - Permissions, filters, and display options * **Chart configurations** - Individual chart settings and interactions * **Filter settings** - Global filters and their default values ## Phase 2: Deployment ### **Step 1: Create guest tokens (Backend)** Guest tokens authenticate embedded users and define their data access permissions. They must be generated securely on your backend service. **Guest token generation:** ```python theme={null} import requests import json def generate_guest_token(api_key, user_info, dashboard_ids, rls_rules): """ Generate a guest token for embedded dashboard access """ payload = { "user": { "username": user_info["username"], "first_name": user_info["first_name"], "last_name": user_info["last_name"] }, "dashboardIds": dashboard_ids, "rls": rls_rules } response = requests.post( "https://api-v1.5x.co/business-intelligence/v1/guest-token", data=json.dumps(payload), headers={"Authorization": f"Bearer {api_key}"} ) return response.json()["data"]["guestToken"] ``` **Guest token parameters:** * **user** (required) - User profile information * **dashboardIds** (required) - List of dashboard IDs the user can access * **rls** (required) - Row-level security rules for data access ### **Step 2: Supply guest tokens to frontend** Create an endpoint in your application that generates and returns guest tokens: ```python theme={null} from flask import Flask, request, jsonify from flask_login import login_required, current_user @app.route("/api/guest-token", methods=["GET"]) @login_required def get_guest_token(): """ Generate guest token for authenticated user """ # Define RLS rules based on user context rls_rules = [ {"clause": f"user_id = '{current_user.id}'"}, {"clause": f"organization_id = '{current_user.organization_id}'"} ] guest_token = generate_guest_token( api_key=os.getenv("5X_API_KEY"), user_info={ "username": current_user.username, "first_name": current_user.first_name, "last_name": current_user.last_name }, dashboard_ids=["1", "2"], # Dashboard IDs user can access rls_rules=rls_rules ) return jsonify({"guestToken": guest_token}) ``` ### **Step 3: Install the Superset Embedded SDK** **From NPM:** ```bash theme={null} npm install --save @superset-ui/embedded-sdk ``` **From CDN:** ```html theme={null} ``` ### **Step 4: Embed the dashboard** **Using ES6 imports:** ```javascript theme={null} import { embedDashboard } from "@superset-ui/embedded-sdk"; // Embed dashboard in your application const dashboard = await embedDashboard({ id: "abc123", // Embedded Dashboard ID from preparation phase supersetDomain: "https://your-workspace.5x.co", mountPoint: document.getElementById("dashboard-container"), fetchGuestToken: () => fetchGuestTokenFromBackend(), dashboardUiConfig: { hideTitle: true, filters: { expanded: true, }, urlParams: { param_name: "value", other_param: "value", }, }, }); ``` **Using CDN:** ```html theme={null} ``` ## Security and access control ### **Row-level security (RLS)** Row-level security provides fine-grained access control by applying additional WHERE clauses to data queries based on user context. **RLS rule examples:** ```json theme={null} { "rls": [ { "clause": "user_id = 'grace_hopper'" }, { "dataset": 16, "clause": "environment = 'production'" }, { "dataset": 42, "clause": "state = 'published'" } ] } ``` **RLS best practices:** * **Dataset-specific rules** - Apply rules to specific datasets when possible * **User context** - Use user information to filter data appropriately * **Security validation** - Never insert untrusted input into RLS clauses * **Testing** - Thoroughly test RLS rules to ensure proper data filtering ### **Domain restrictions** **Configure allowed domains:** * **Exact domain matching** - Specify exact domains (no wildcards) * **Protocol inclusion** - Include https\:// or http\:// in domain specifications * **Environment coverage** - Include all relevant environments (dev, staging, prod) * **Security review** - Regularly review and update allowed domains ### **Authentication flow** **Guest token lifecycle:** * **1-hour expiration** - Guest tokens expire after 1 hour * **Automatic refresh** - SDK automatically refreshes tokens * **Secure generation** - Tokens generated only on trusted backend * **No session persistence** - Embedded dashboards don't maintain 5X sessions ## Best practices ### **Security best practices** **Backend only** Always generate guest tokens on your backend service, never expose API keys to frontend. **Minimal access** Grant only the minimum permissions necessary for each user's role and needs. **Ongoing monitoring** Regularly review and audit access permissions and RLS rules. **Validate inputs** Always validate and sanitize user inputs before using them in RLS clauses. ## Troubleshooting ### **Common embedding issues** **Possible causes:** * Invalid or expired guest tokens * Incorrect API key configuration * Missing user information in token * Network connectivity issues **Solutions:** * Verify API key and token generation * Check user information completeness * Test network connectivity * Review token expiration handling **Possible causes:** * RLS rules too restrictive * Missing dashboard permissions * Data source connection issues * Query errors or timeouts **Solutions:** * Review and adjust RLS rules * Verify dashboard access permissions * Check data source connections * Test queries independently **Possible causes:** * Incorrect iframe sizing * CSS conflicts with parent application * Responsive design problems * Browser compatibility issues **Solutions:** * Adjust iframe dimensions and styling * Resolve CSS conflicts * Test responsive design * Check browser compatibility *** Continue your Business Intelligence journey with these related guides. **Create dashboards** Build dashboards that can be embedded in your applications. # Datasets & data exploration Source: https://documentation.5x.co/core-features/business-intelligence/data-exploration Explore datasets, write custom SQL queries, and understand your data structure with powerful data exploration tools **Focus:** Learn how to explore your data effectively, write custom SQL queries, and understand data quality and structure to build better visualizations. ## Data exploration overview Data exploration is the foundation of effective analytics. Before creating charts and dashboards, you need to understand your data - its structure, quality, relationships, and patterns. 5X Business Intelligence provides powerful tools for exploring datasets and writing custom SQL queries. ### **Why data exploration matters** **Ensure accuracy** Understanding data quality helps you build reliable visualizations and avoid misleading insights. **Find insights** Explore data to discover patterns, trends, and relationships that inform your analysis. **Improve performance** Understanding data structure helps you write efficient queries and optimize chart performance. **Better analysis** Thorough data exploration leads to more accurate analysis and better business decisions. ## Dataset management ### **What are datasets?** Datasets in 5X Business Intelligence are logical representations of your data sources that serve as the foundation for all analytics and visualizations. They act as a bridge between your raw data warehouse tables and the charts, dashboards, and reports you create. **Key characteristics of datasets:** * **Logical abstraction** - Represent tables, views, or custom SQL queries from your data warehouse * **Metadata rich** - Include schema information, column types, relationships, and business context * **Performance optimized** - Include caching and query optimization features * **Business focused** - Organized and documented for business users **Types of datasets:** * **Physical tables** - Direct representation of warehouse tables * **Views** - Pre-defined SQL views from your data warehouse * **Custom SQL** - Virtual datasets created from custom SQL queries * **Metrics layer datasets** - Business-friendly views from 5X's metrics layer ### **Creating datasets** **From App Connections:** 1. **Navigate to datasets** - Go to the datasets section in Business Intelligence 2. **Select data source** - Choose from your configured App Connections 3. **Browse tables** - Explore available tables and views 4. **Create dataset** - Click "Create Dataset" for your selected table 5. **Configure settings** - Set name, description, and basic properties 5X Business Intelligence Dataset Creation Flow **From SQL Lab:** 1. **Access SQL Lab** - Navigate to SQL Lab from the Business Intelligence menu 5X Business Intelligence Dataset Creation Flow from SQL Lab 2. **Write custom query** - Create SQL query in SQL Lab 3. **Test and validate** - Execute query to verify results 4. **Save as dataset** - Use "Save as Dataset" option 5. **Configure metadata** - Add business context and documentation 5X Business Intelligence Dataset Creation Flow from SQL Lab ### **Dataset best practices** **Descriptive names** Use consistent, descriptive naming conventions that reflect business purpose and data source. **Business context** Include detailed descriptions, business rules, and usage guidelines for each dataset. **Organization** Use tags and categories to organize datasets for easy discovery and management. **Access control** Implement appropriate security policies and row-level security where needed. ## SQL Lab ### **Introduction to SQL Lab** SQL Lab is 5X Business Intelligence's powerful query interface that allows you to write and execute custom SQL queries directly against your data warehouse. **Key features:** * **Full SQL support** - Write complex queries with joins, subqueries, and advanced functions * **Query history** - Track and reuse previous queries * **Query sharing** - Share queries with team members * **Result export** - Export query results in various formats * **Query optimization** - Built-in suggestions for query performance ### **Writing effective queries** **Query structure best practices:** ```sql theme={null} -- Example: Well-structured query with clear formatting SELECT DATE_TRUNC('month', order_date) AS month, region, COUNT(*) AS order_count, SUM(order_amount) AS total_revenue, AVG(order_amount) AS avg_order_value FROM orders WHERE order_date >= '2024-01-01' AND order_status = 'completed' GROUP BY 1, 2 ORDER BY month DESC, total_revenue DESC LIMIT 100; ``` **Query optimization tips:** * **Use appropriate filters** - Limit data with WHERE clauses * **Select only needed columns** - Avoid SELECT \* for better performance * **Use proper indexing** - Structure queries to leverage database indexes * **Limit result sets** - Use LIMIT to control output size ### **Advanced SQL features** **Window functions:** ```sql theme={null} -- Example: Using window functions for advanced analytics SELECT customer_id, order_date, order_amount, SUM(order_amount) OVER ( PARTITION BY customer_id ORDER BY order_date ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW ) AS running_total, ROW_NUMBER() OVER ( PARTITION BY customer_id ORDER BY order_date DESC ) AS order_rank FROM orders; ``` **Common table expressions (CTEs):** ```sql theme={null} -- Example: Using CTEs for complex analysis WITH monthly_sales AS ( SELECT DATE_TRUNC('month', order_date) AS month, SUM(order_amount) AS total_sales FROM orders GROUP BY 1 ), sales_growth AS ( SELECT month, total_sales, LAG(total_sales) OVER (ORDER BY month) AS prev_month_sales, (total_sales - LAG(total_sales) OVER (ORDER BY month)) / LAG(total_sales) OVER (ORDER BY month) * 100 AS growth_rate FROM monthly_sales ) SELECT * FROM sales_growth; ``` ## Troubleshooting ### **Common exploration issues** **Possible causes:** * Insufficient permissions for dataset access * App connection configuration problems * Data source connectivity issues * Row-level security restrictions **Solutions:** * Verify user permissions and roles * Check App Connection settings * Test data source connectivity * Review security policies **Possible causes:** * Inefficient query structure * Large dataset volumes * Missing database indexes * Complex calculations or joins **Solutions:** * Optimize query structure and logic * Use appropriate filters and limits * Check database indexing strategy * Consider data aggregation approaches *** **Connect to data** Learn how to set up App Connections to access your data sources. **Create visualizations** Use your data exploration insights to build compelling charts and dashboards. # Overview Source: https://documentation.5x.co/core-features/business-intelligence/overview Interactive dashboards and self-service analytics - create compelling visualizations for your stakeholders **Focus:** Learn how to create interactive dashboards, build compelling visualizations, and enable self-service analytics for your organization using 5X Business Intelligence. ## Overview 5X Business Intelligence delivers powerful self-service analytics capabilities that enable your team to create interactive dashboards, build compelling visualizations, and explore data insights without technical barriers. Enhanced with 5X's integrated platform features, it provides everything you need to transform raw data into actionable business insights. **Powered by Apache Superset** 5X Business Intelligence is built on Apache Superset, the leading open-source business intelligence platform. We've enhanced it with 5X-specific integrations and streamlined the experience for our integrated data platform ecosystem. 5X Business Intelligence Dashboard Interface ## Key capabilities ### **Interactive dashboard creation** * **Drag-and-drop interface** for building compelling dashboards * **Rich visualization library** with 50+ chart types * **Real-time data updates** with automatic refresh capabilities * **Responsive design** that works across all devices * **Collaborative features** for team-based dashboard development ### **Self-service analytics** * **No-code chart building** for business users * **Intuitive data exploration** with guided workflows * **Custom SQL support** for advanced users * **Data profiling** to understand data quality and structure * **Cross-filtering** and interactive drill-down capabilities ### **Integrated data connectivity** * **App Connections integration** - Connect to your data warehouse seamlessly * **Metrics Store support** - Access unified business metrics * **Real-time data access** - Query your warehouse directly * **Multiple data sources** - Support for Snowflake, BigQuery, Redshift, PostgreSQL ### **Advanced features** * **Email alerts and reports** - Automated dashboard delivery * **Dashboard embedding** - Embed analytics in external applications * **Row-level security** - Fine-grained access control * **Custom branding** - Match your organization's visual identity ## Getting started Ready to create your first dashboard? Here's how to get started: ### **Step 1: Set up your BI data connection** Before you can create dashboards, you need to configure a Business Intelligence app connection. This is essential because: * **5X BI requires a dedicated connection** - BI needs its own app connection with specific permissions * **Choose your data source** - Connect directly to your warehouse or use 5X's Metrics Store for unified metrics * **Secure access** - BI connections use read-only permissions to protect your data while enabling analytics Learn how to create a Business Intelligence app connection to access your data warehouse or metrics layer. ## Core documentation Explore the comprehensive guides and tutorials to master 5X Business Intelligence: **Build interactive dashboards** Learn how to create compelling dashboards with charts, filters, and interactive elements. **Create visualizations** Master the art of building compelling charts and visualizations from your data. **Explore your data** Learn how to explore datasets, write custom SQL, and understand your data structure. **Connect to data sources** Configure Business Intelligence app connections to access your data warehouse. **Embed in applications** Integrate dashboards into your applications with secure embedding capabilities. ### **Streamlined administration** 5X Business Intelligence removes complexity by: * **Hiding database management** from the BI interface * **Integrating user management** with 5X platform users * **Simplifying deployment** with managed infrastructure * **Providing unified support** across all platform capabilities *** **Start with quickstart** Get your workspace running and see Business Intelligence in action in 15 minutes. # Chats types Source: https://documentation.5x.co/core-features/conversational-ai/chats Manage, share, and organize your conversational AI chats with team collaboration features ## Share a chat with teammates Click the Lock/Share icon near the send button. In Sharing Settings, choose: * Only me, or * Specific users (auto-complete from users who have access to your 5X workspace). Save — the chat appears in Team Chats for those users. Settings page with app connections highlighted in sidebar ## Flagged Chats Use Flagged Chats to mark conversations for review (accuracy checks, follow-ups, or audit). * Flag from the chat card. * Open Flagged Chats to see the list, token usage, shared users, and quickly revisit items. Settings page with app connections highlighted in sidebar ## Chat Settings (Tone, Domain, Prompts) Customise how the AI responds: * **Tone:** Fact-based, Expressive, or Balanced * **Domain:** Sales, Marketing, Finance, Leadership, etc. * **Response Style:** Concise vs. descriptive * **Additional Prompting:** Add custom instructions to guide responses across the session Settings page with app connections highlighted in sidebar # Overview Source: https://documentation.5x.co/core-features/conversational-ai/overview Natural language interface for querying data and generating insights from your Metrics Store **Focus:** Learn how 5X Conversational AI makes data accessible to everyone through natural language conversations, enabling users to get insights without writing SQL or building dashboards. ## What is Conversational AI? The Conversational AI feature lets business and data users ask questions like: * "Top 10 countries by high-value customers and their average churn score?" * "What's revenue growth QoQ by region?" It discovers relevant metrics from your Metrics Store and automatically runs the appropriate warehouse queries. ## Access & Navigation Open the left navigation bar in the 5X Platform. Click Conversational AI. You'll land on New Chat, with tabs for My Chats, Team Chats, and Flagged Chats. Settings page with app connections highlighted in sidebar ## Start a New Chat * Type your question in English in the input box. * Press Enter to run the query. * Typical response time: \~10–15 seconds (may be faster when cached). ## Project Selection (Metrics Store) Choose which Metrics project(s) to query: * **All Projects** (recommended if you're unsure which project has the right metric) * The system matches your question to the best project before execution. * **Specific Project** (faster when you know exactly where the metrics live) If you know the project (e.g., Sales Metrics), select it to reduce cross-project checks and speed up response. Settings page with app connections highlighted in sidebar ## Voice Input (Mic) Click the mic icon to speak your question instead of typing. Make sure your browser permissions allow microphone access. ## Sync Now (Metadata Refresh) If a developer just added/updated metrics in the Metrics Store and you don't see them yet: * Click Sync Now to force a metadata refresh. * Conversational AI ingests the latest metadata within \~5 minutes. # Query, Tables, Charts Source: https://documentation.5x.co/core-features/conversational-ai/querytablescharts Explore query results, visualizations, and manage LLM settings ## Results: Query, Tables, Charts, Exports Each response supports: * **Table View / Chart View toggles** — Switch between data table and visualization * **Expand for a full-screen view** — View results in a larger format * **Download as PNG** — Export charts for sharing and decks * **SQL** — View the query used to fetch the data from the warehouse Settings page with app connections highlighted in sidebar ## Related Questions & Metadata Under each answer: * **Related Questions:** Contextual follow-ups to keep exploring. * **Query Metadata:** Token usage, runtime ## LLM / API Key Settings You can use the default 5X Managed model or your own API key: * **Providers supported** (examples): Google Gemini, OpenAI, Anthropic * **5X surfaces the latest stable versions** compatible with Conversational AI * **When you add your key**, subsequent queries run via your provider's billing Settings page with app connections highlighted in sidebar ## Troubleshooting Run again with All Projects, or explicitly select the correct project. Click Sync Now and re-run in \~5 minutes. Narrow the question (date range, fewer dimensions), or target a specific project. * Open the metric definition in the semantic layer, confirm filters and time grain. * Re-ask with domain hints (e.g., "Sales – net revenue, last 90 days, by region"). Prefer specific projects, concise questions, and related follow-ups (reduces re-grounding). ## Privacy & Governance * **Access controls** follow your 5X workspace permissions. * **Share only with permitted users** via Team Chats. * **API keys** (when provided) are workspace-scoped and used only for your queries. ## Glossary * **Semantic Layer / Metric Store:** Central definitions of metrics and dimensions. * **Sync Now:** Forces a metadata refresh for newly added/edited metrics. * **Related Questions:** Auto-suggested next steps based on the current result and ingested metadata. * **Tokens:** LLM processing units; higher tokens ≈ longer inputs/outputs. # Creating Data & AI Apps Source: https://documentation.5x.co/core-features/data-ai-apps/app-creation Learn how to create new Streamlit applications using 5X managed repositories or import existing ones from GitHub **Focus:** Master the app creation process, understand repository options, and set up your first Data & AI App for development. Creating Data & AI Apps in 5X is straightforward and flexible. You can either start with a fresh Streamlit application using 5X's managed repository, or import an existing Streamlit app from your GitHub repository. ## App creation workflow ### **Step 1: Access Manage Apps** Navigate to the **Data & AI Apps** section in your 5X platform and click on **Manage Apps** to access the application management interface. Manage Apps Interface The Manage Apps interface shows all your applications with their current status, deployment options, and management controls. ### **Step 2: Create new app** Click the **+ Create new app** button to open the app creation modal. ## Repository options Choose how you want to create your Streamlit application: Choose **5X Managed** when you want to create a new Streamlit application from scratch: 5X Managed Repository Option **When to use 5X Managed:** * Creating a new application from scratch * Learning Streamlit development * Quick prototyping and experimentation * Applications that don't require complex Git workflows **What you get:** * Pre-configured Streamlit environment * Basic app structure and dependencies * Ready-to-use development environment * Automatic repository management Choose **GitHub** when you have an existing Streamlit application or want to maintain your own repository: GitHub Repository Option **When to use GitHub:** * Importing existing Streamlit applications * Maintaining version control with your team * Complex applications with custom dependencies * Applications that need to be shared across multiple environments **Requirements:** * Public or accessible private GitHub repository * Repository must contain a valid Streamlit application * Proper `requirements.txt` file for dependencies ## Next steps After creating your app, you're ready to start development: **Start developing** Learn how to develop and test your application using the 5X IDE. **Deploy your app** Understand how to publish and configure deployment options for your application. ## Troubleshooting ### **Common issues** **Issue:** Cannot access GitHub repository **Solutions:** * Verify the repository URL is correct * Ensure the repository is public or you have proper access permissions * Check that the repository contains a valid Streamlit application * Confirm the repository has a `requirements.txt` file **Issue:** App creation fails due to repository structure **Solutions:** * Ensure your repository has a main Streamlit file (usually `app.py` or `main.py`) * Verify the `requirements.txt` file includes Streamlit and other dependencies * Check that the repository structure follows Streamlit conventions * Make sure there are no conflicting files or configurations **Issue:** App configuration validation fails **Solutions:** * Use valid characters in repository and app names * Ensure app names are descriptive and unique * Verify emoji selection is working properly * Check that all required fields are filled out *** **Develop your app** Learn how to build and test your Streamlit application using the 5X IDE. **IDE configuration** Set up your development environment for efficient app development. # Overview Source: https://documentation.5x.co/core-features/data-ai-apps/overview Build and deploy custom data-driven applications with embedded analytics and AI capabilities **Focus:** Create powerful, interactive applications that leverage your data warehouse directly, enabling you to build custom solutions for specific business needs. 5X Data & AI Apps are custom applications powered by Streamlit that connect directly to your data warehouse, enabling you to build interactive dashboards, data tools, and AI-powered solutions tailored to your specific business requirements. **Streamlit-Powered** All 5X Data & AI Apps are built on the Streamlit framework, providing a powerful and flexible foundation for creating interactive data applications. ### **Core capabilities** **Powerful app development** Build sophisticated applications using the industry-standard Streamlit framework with full Python capabilities. **Real-time data connectivity** Connect directly to your data warehouse for real-time data access and processing. **Multiple deployment options** Choose from on-demand, always-on, or scheduled deployment based on your use case requirements. **IDE integration** Develop and test your applications using the 5X IDE with full VS Code functionality. **Secure access control** Implement fine-grained access control with role-based permissions for different user groups. **Seamless ecosystem** Leverage the full 5X platform including metrics layer, orchestration, and business intelligence. ## The future of data applications ### **Building on your data foundation** Data & AI Apps represent the evolution of how organizations interact with their data. Instead of building applications separately and then connecting to data sources, you can now build applications directly on top of your data warehouse. **Why this matters:** * **Real-time insights** - Applications always work with the freshest data * **Consistent data models** - All apps use the same underlying data definitions * **Reduced complexity** - No need to manage separate data pipelines for applications * **Better performance** - Direct warehouse access eliminates data transfer overhead * **Unified governance** - Single point of control for data access and security ### **Platform integration benefits** **Unified metrics** Access consistent business metrics and definitions across all your applications. **Automated workflows** Integrate apps into your data orchestration workflows for automated updates and processing. **Enhanced analytics** Combine app functionality with traditional BI capabilities for comprehensive insights. **AI-powered features** Integrate natural language processing and AI capabilities into your applications. ## Getting started Ready to create your first Data & AI App? Here's how to get started: ### **Step 1: Access the Manage Apps section** Navigate to the **Data & AI Apps** section in your 5X platform and click on **Manage Apps** to access the application management interface. Manage Apps Interface ### **Step 2: Create your first app** Click **+ Create new app** to start building your application. You'll have two options: * **5X Managed** - Create a new Streamlit app from scratch using 5X's managed repository * **GitHub** - Import an existing Streamlit application from your GitHub repository Create New App Modal ## Core documentation Explore the comprehensive guides to master Data & AI Apps development: **Build your first app** Learn how to create new Streamlit applications using 5X managed repositories or import existing ones. **Deploy your apps** Learn how to publish applications and configure deployment options for different use cases. ## Key benefits ### **Rapid application development** Build sophisticated data applications in hours, not weeks. The combination of Streamlit's rapid development capabilities and 5X's integrated platform eliminates the complexity of setting up development environments and managing data connections. ### **Enterprise-grade deployment** Deploy applications with confidence using enterprise-grade infrastructure: * **Scalable hosting** - Automatic scaling based on demand * **Security controls** - Role-based access and data protection * **Monitoring** - Built-in performance monitoring and alerting * **Compliance** - SOC 2, GDPR, and HIPAA compliance support ### **Cost-effective operations** Reduce operational overhead with managed infrastructure: * **No server management** - Focus on building applications, not managing infrastructure * **Automatic updates** - Platform updates and security patches handled automatically * **Resource optimization** - Efficient resource usage with on-demand scaling * **Unified billing** - Single platform billing instead of multiple tool subscriptions ## Getting help **Need assistance with your first Data & AI App?** Our team is here to help you get started. Reach out to us via: * **Email:** [support@5x.co](mailto:support@5x.co) * **Intercom:** Available in your 5X platform * **Customer Success Manager:** Your dedicated CSM for personalized assistance We can provide demos, help set up your first use case, and guide you through the development process. *** # Publishing & Deployment Source: https://documentation.5x.co/core-features/data-ai-apps/publishing Publish your applications and configure deployment options for different use cases and environments **Focus:** Learn how to publish your Data & AI Apps, configure deployment options, and manage access controls for different user groups. Publishing your Data & AI App makes it available to users and configures how it runs in production. You can choose from multiple deployment options and set up access controls to ensure your application is available to the right users at the right time. ## Publishing workflow ### **Access the publish dialog** To publish your app, navigate to the **Manage Apps** section and click the three-dots menu next to your application, then select **Publish**: Publish Menu Option This opens the **Publish app** dialog where you can configure all deployment settings: Publish App Dialog ## App configuration ### **Basic app settings** Configure the essential settings for your published application: **App emoji:** * Select an emoji that represents your application * This appears in the navigation and app listings * Choose something that users will easily recognize **App name:** * Enter the display name for your application * This is what users will see in the navigation * Use clear, descriptive names that indicate the app's purpose **Branch:** * Select which Git branch to deploy from * Typically `main` or `master` for production deployments * You can use different branches for different environments ### **Technical configuration** **Python version:** * Select the Python version for your application * Choose based on your app's requirements and dependencies * Common options include Python 3.8.20, 3.9.23, 3.10.18, 3.11.13, 3.12.11, and 3.13.4 **Compute profile:** * Configure the compute resources for your application * Options range from basic to high-performance configurations * Choose based on your app's computational needs and expected usage * **Custom profiles**: Create and manage custom compute profiles from **Settings → Compute Profiles** in the platform Compute Profile Configuration ## Deployment options ### **Understanding deployment types** Choose the deployment option that best fits your use case: Deployment Options **Best for:** Production applications with consistent usage **Characteristics:** * Application runs continuously * Immediate access for users * Higher resource costs but better user experience * Suitable for mission-critical applications **Use cases:** * Production dashboards and reports * Customer-facing applications * Real-time monitoring tools * Applications requiring immediate availability **Best for:** Applications with predictable usage patterns **Characteristics:** * App runs only during specified hours * Outside scheduled hours, starts on-demand * Balances cost and availability * Configurable time zones and schedules **Configuration options:** * **Time range:** Set start/end time or use CRON Expression * **Days:** Choose weekdays only or full week * **Timezone:** Configure for your location * **Automatic timezone:** Detect user's location **Use cases:** * Business hours applications * Regional applications with specific time zones * Cost-optimized production deployments * Applications with predictable usage patterns **Best for:** Development, testing, and low-frequency usage **Characteristics:** * App starts up when a user tries to access it * It usually takes some time to start * Automatically shuts down after 30 minutes of inactivity * Cost-effective for applications with sporadic usage **Use cases:** * Development and testing environments * Internal tools with occasional usage * Prototype applications * Applications with unpredictable usage patterns ## Access control and permissions ### **Role-based access control** Configure who can access your application using role-based permissions: Access Control Configuration ### **Permission best practices** **Minimal access** Grant only the minimum permissions necessary for users to perform their tasks. **Access audits** Regularly review and update access permissions as roles change. **Permission tracking** Document who has access to what and why for compliance purposes. **Verify permissions** Test access controls to ensure they work as expected. ## Publishing process ### **Step-by-step publishing** Set your app name, emoji, branch, and environment Choose appropriate compute resources for your application Select on-demand, always-on, or scheduled deployment Set up view and edit permissions for different user roles Review all settings and click Publish to deploy your application ### **Post-publishing** After publishing, your application will: 1. **Appear in navigation** - Users with view permissions will see the app in their left navigation 2. **Be accessible** - The app will be available based on your deployment configuration 3. **Show in Manage Apps** - You can monitor and manage the app from the Manage Apps interface 4. **Display status** - The app status will show as "Published" in the apps list ## Managing published apps ### **App status management** **Published apps:** * Visible to users with appropriate permissions * Accessible based on deployment configuration * Can be managed through the Manage Apps interface **Unpublished apps:** * Not visible to regular users * Still accessible to Admin users and users with Global edit access * Can be republished at any time ### **Edit app details** You can modify app settings after publishing by clicking **Edit details** from the three-dots menu. This opens the same configuration dialog as publishing, allowing you to: * Update app settings and configuration * Change deployment options * Modify access permissions * Update compute profiles ### **Unpublishing apps** To unpublish an app: 1. Go to the Manage Apps interface 2. Click the three-dots menu next to your app 3. Select **Unpublish** (this option appears for published apps) 4. Confirm the unpublishing action Unpublished apps remain accessible to Admin users and users with Global edit access, but are hidden from regular users. ## Troubleshooting ### **Common publishing issues** **Issue:** App fails to publish **Solutions:** * Check that all required fields are filled out * Verify the selected branch exists and contains valid code * Ensure compute profile is appropriate for your app * Check for any configuration errors or conflicts **Issue:** App takes too long to start **Solutions:** * Consider upgrading to a higher compute profile * Optimize your app's startup time and dependencies * Use always-on deployment for frequently accessed apps * Review and optimize your requirements.txt **Issue:** Users cannot access the published app **Solutions:** * Verify user roles and permissions are correctly configured * Check that users have the appropriate view permissions * Ensure the app is published and not unpublished * Verify deployment configuration is correct *** # Development Workflows Source: https://documentation.5x.co/core-features/ide/development-workflows Learn how to develop dbt models, Python applications, Streamlit dashboards, and Metrics project using the IDE with best practices and examples **Focus:** Master the development workflows for different types of data applications, from dbt transformations to Python scripts and interactive Streamlit dashboards. The 5X IDE provides specialized development environments for the most common data development tasks. This guide covers practical workflows for building data models, applications, and dashboards with real examples and best practices. **Development Environment** The IDE comes pre-configured with multiple Python versions, dbt environments, and development tools. No additional setup required for basic development tasks. ## Python development ### **Python environment overview** The IDE comes pre-installed with multiple Python versions managed through pyenv, providing flexibility for different project requirements and dependency compatibility. **Available Python versions:** * **Python 3.8.20** - Extended legacy support for older projects * **Python 3.9.23** - Legacy support for older projects * **Python 3.10.18** - Stable version with good package compatibility * **Python 3.11.13** - Default version (set by PYENV\_VERSION) * **Python 3.12.11** - Latest stable with performance improvements * **Python 3.13.4** - Cutting-edge features and optimizations **View installed versions:** ```bash theme={null} ls /root/.pyenv/versions ``` ### **Virtual environment management** Python virtual environments provide isolated dependency management for your projects, preventing conflicts between different project requirements. **Create a virtual environment:** ```bash theme={null} # Using Python 3.11.13 (default) /root/.pyenv/versions/3.11.13/bin/python -m venv my_project_env # Using specific Python version /root/.pyenv/versions/3.10.18/bin/python -m venv legacy_project_env ``` **Activate and manage environments:** ```bash theme={null} # Activate environment source my_project_env/bin/activate # Verify active environment (should show your env path) which python # Deactivate when finished deactivate ``` ### **Dependency management best practices** Maintain project dependencies using requirements.txt files for reproducible environments across team members and deployment targets. **Create requirements.txt:** ```text theme={null} # Core data processing pandas==2.0.3 numpy==1.24.3 # API and web requests requests==2.31.0 urllib3==2.0.4 # Visualization matplotlib==3.7.2 seaborn==0.12.2 # Development tools jupyter==1.0.0 pytest==7.4.0 ``` **Install and manage dependencies:** ```bash theme={null} # Activate environment first source my_project_env/bin/activate # Install from requirements file pip install -r requirements.txt # Install additional packages and update requirements pip install scikit-learn==1.3.0 pip freeze > requirements.txt ``` ### **Python development examples** **Data processing script:** ```python theme={null} import pandas as pd import numpy as np from sqlalchemy import create_engine def process_customer_data(connection_string): """Process customer data from warehouse""" engine = create_engine(connection_string) # Load data df = pd.read_sql("SELECT * FROM customers", engine) # Data transformations df['full_name'] = df['first_name'] + ' ' + df['last_name'] df['customer_tier'] = pd.cut(df['total_spent'], bins=[0, 100, 500, 1000, float('inf')], labels=['Bronze', 'Silver', 'Gold', 'Platinum']) # Save processed data df.to_sql('processed_customers', engine, if_exists='replace', index=False) return df if __name__ == "__main__": # Your connection string here conn_str = "your_connection_string" result = process_customer_data(conn_str) print(f"Processed {len(result)} customers") ``` **API integration example:** ```python theme={null} import requests import json from datetime import datetime class DataAPI: def __init__(self, base_url, api_key): self.base_url = base_url self.headers = {'Authorization': f'Bearer {api_key}'} def fetch_data(self, endpoint, params=None): """Fetch data from API endpoint""" response = requests.get( f"{self.base_url}/{endpoint}", headers=self.headers, params=params ) response.raise_for_status() return response.json() def process_and_save(self, endpoint, db_connection): """Fetch, process, and save data to database""" data = self.fetch_data(endpoint) # Process data df = pd.DataFrame(data) df['processed_at'] = datetime.now() # Save to database df.to_sql('api_data', db_connection, if_exists='append', index=False) return df ``` ## dbt development ### **dbt Power User extension (recommended approach)** The dbt Power User extension provides the most integrated development experience, automatically using your configured dbt settings from Settings → Credentials including version selection, database connections, and target configuration. dbt Power User Extension **Key workflows:** **Run and test models** Execute individual models, selections, or entire dbt projects with integrated test runner **Understand dependencies** Interactive dependency graphs showing upstream and downstream model relationships **Generate docs** Create and view dbt documentation with integrated preview and automatic refresh **Preview compiled SQL** See the actual SQL that will be executed before running models ### **Command-line dbt development** For users preferring terminal-based workflows, the IDE provides pre-configured dbt virtual environments for each supported version. **Activate dbt environment:** ```bash theme={null} # List available dbt environments ls /root/.venv # Activate specific dbt version source /root/.venv/dbt-1.8.9/bin/activate # Verify dbt installation dbt --version ``` **Available dbt versions:** * **dbt-1.6.18** (`/root/.venv/dbt-1.6.18/`) - Legacy support * **dbt-1.7.19** (`/root/.venv/dbt-1.7.19/`) - Stable version * **dbt-1.8.9** (`/root/.venv/dbt-1.8.9/`) - Current stable * **dbt-1.9.10** (`/root/.venv/dbt-1.9.10/`) - Latest features ### **dbt development workflow** **Common dbt commands:** ```bash theme={null} # Navigate to your dbt project directory cd /path/to/your/dbt/project # Run entire project dbt run # Run specific models dbt run --select staging.stg_customers+ # Test your models dbt test # Generate documentation dbt docs generate dbt docs serve ``` **Model development example:** ```sql theme={null} -- models/staging/stg_customers.sql SELECT customer_id::int AS customer_id, LOWER(TRIM(email)) AS email, INITCAP(first_name) AS first_name, INITCAP(last_name) AS last_name, created_at::timestamp AS created_at, CASE WHEN status = 'A' THEN 'active' WHEN status = 'I' THEN 'inactive' ELSE 'unknown' END AS status FROM {{ source('crm', 'customers') }} WHERE customer_id IS NOT NULL ``` **Model testing:** ```yaml theme={null} # models/schema.yml version: 2 models: - name: stg_customers description: "Cleaned customer data from CRM system" columns: - name: customer_id description: "Unique customer identifier" tests: - unique - not_null - name: email description: "Customer email address" tests: - not_null - unique ``` ### **Lineage visualization** The IDE provides powerful lineage visualization capabilities that help you understand data flow and model dependencies throughout your dbt project. dbt Model Lineage Visualization **To view lineage:** 1. Open any dbt model file in the editor 2. Navigate to the Lineage tab in the IDE interface 3. Explore interactive dependency graphs showing: * Upstream models and sources feeding into current model * Downstream models consuming current model output * Cross-project dependencies and external table references **Lineage features:** * **Interactive navigation** - Click nodes to jump between related models * **Dependency depth control** - Adjust how many levels of dependencies to display * **Impact analysis** - Understand which models will be affected by changes * **Visual debugging** - Identify circular dependencies and optimization opportunities ### **Running dbt commands** Execute dbt commands directly from the IDE for any valid dbt project without leaving your development workspace. **To run dbt commands:** Open a dbt project folder in the IDE Click on the icon from the top-right status bar ```bash theme={null} # List available dbt environments ls /root/.venv # Activate specific dbt version source /root/.venv/dbt-1.8.9/bin/activate # Verify dbt installation dbt --version ``` A command input box will appear, allowing you to enter the desired dbt command Confirm the command to execute it in the terminal **Terminal session management:** If similar commands were executed previously for the same project, the IDE will prompt you to either: * **Open a new terminal session** - Start a fresh terminal for the command * **Continue using existing terminal** - Reuse the current terminal session This integrated feature simplifies the dbt workflow, enabling you to build, test, and manage transformations without leaving your development workspace. **Example dbt commands:** ```bash theme={null} # Run all models dbt run # Run specific models dbt run --select staging.stg_customers+ # Test models dbt test # Generate documentation dbt docs generate # Compile models dbt compile ``` ## Cube development ### **Cube creation** Create new cubes directly from the IDE without leaving your development environment, providing a seamless and integrated experience for managing cube creation. **To create a cube:** Open any file within the cubes repository Click on the icon located on the top-right corner of the status bar The Cube Server will automatically start, enabling the cube creation process A new file tab will open, displaying a list of available schemas. Select the desired schema Proceed to define and create cubes as needed This workflow provides a seamless and integrated experience for managing cube creation within the IDE, eliminating the need to switch between different tools or environments. ### **Cube server start** Start the Cube Server for a specific active file tab directly from the IDE toolbar, giving you full control over active Cube Server sessions. **To start the Cube Server:** Open any file within your cubes repository Click on the icon from the IDE toolbar Once initiated, the Cube Server will launch and can be accessed locally via `http://localhost:4000` **Server conflict management:** If there is an existing Cube Server instance running, a prompt will appear asking you to: * **Stop the currently active server and start a new one** - Terminate the existing instance and launch a fresh server * **Cancel to retain the current session** - Keep the existing server running This ensures that multiple servers do not conflict and that you retain full control over active Cube Server sessions. ## Running Streamlit applications and Python files ### **Running Streamlit applications** Run Streamlit applications directly from the IDE with automatic environment setup and dependency management. **To run a Streamlit application:** Open the `streamlit_app.py` file from the Streamlit repository Click on the icon in the IDE toolbar The IDE will prompt you to select the desired Python version for execution Upon confirmation, a virtual environment will be created automatically All dependencies listed in the `requirements.txt` file will be installed within the environment Once setup is complete, the Streamlit application will launch successfully ### **Running Python files** Execute standalone Python scripts with the same streamlined workflow as Streamlit applications. **To run a Python file:** Open any Python file (`.py`) in the IDE Click on the icon in the IDE toolbar Choose the desired Python version for execution A virtual environment will be created automatically if needed Dependencies from `requirements.txt` will be installed automatically The Python file will run successfully with output displayed in the terminal **Benefits of integrated execution:** * **Environment consistency** - Automatic virtual environment creation ensures consistent execution environments * **Dependency management** - Automatic installation of dependencies from `requirements.txt` reduces manual setup overhead * **Version selection** - Choose the appropriate Python version for your project requirements * **Seamless workflow** - Run applications and scripts without leaving the IDE or switching contexts This process ensures environment consistency and reduces manual setup overhead for Streamlit or Python-based workflows, allowing you to focus on development rather than environment configuration. # IDE Interface Source: https://documentation.5x.co/core-features/ide/interface Navigate the IDE interface, understand core components, and master the development environment for efficient data development **Focus:** Learn how to navigate and use the 5X IDE interface effectively, including VS Code features, database explorer, source control, and session management. The 5X IDE provides a full-featured VS Code experience directly in your browser, with specialized tools for data development. This guide covers all interface components and how to use them effectively for your data projects. **VS Code Familiarity** If you're familiar with VS Code, you'll feel right at home. The IDE includes all standard VS Code features plus specialized data development tools. ## Starting IDE sessions ### **Session initialization** Starting your IDE session is straightforward and takes some time to fully load: 5X IDE Interface Overview Go to the **IDE** tab in your 5X platform navigation Click **Start Session** to initialize your development environment Allow some time for the environment to fully load Access your repositories and start building in the familiar VS Code interface ### **Session lifecycle management** Understanding session behavior helps you work efficiently: **Session characteristics:** * **Auto-termination**: Sessions end after 30 minutes of inactivity * **Persistent state**: Environment configurations and saved changes persist between sessions * **Manual restart**: Use IDE controls to restart or reset when needed * **Resource management**: Sessions are optimized for data development workloads ## Core interface components ### **File explorer** The File Explorer displays all repositories created through the 5X platform, providing intuitive navigation and file management capabilities. 5X IDE File Explorer **Key features:** * **Repository navigation** - Browse file trees and project structures across all integrated repositories * **File operations** - Create, edit, delete, and organize files with standard operations * **Project documentation** - Access README files, configuration files, and project documentation * **Drag-and-drop support** - Organize files and folders within projects * **Multi-repository view** - Work with multiple projects simultaneously **File explorer benefits:** * **Integrated repositories** - All your 5X-managed repositories appear automatically * **Version control integration** - See Git status indicators for files and folders * **Search functionality** - Find files quickly across all repositories * **Context menus** - Right-click for file operations and Git commands ### **Source control integration** Comprehensive Git workflow management directly within the IDE environment: IDE Source Control Interface **Available operations:** * **Commit changes** with descriptive commit messages and staged/unstaged file management * **Branch management** - Create, switch, and merge branches with visual branch management * **Remote operations** - Pull and push changes to remote repositories * **Change tracking** - View file changes, diffs, and commit history with side-by-side comparison * **Conflict resolution** - Resolve merge conflicts with visual merge tools and guided resolution **Git workflow features:** * **Visual diff viewer** - See exactly what changed in each file * **Commit history** - Browse commit history with detailed commit information * **Branch visualization** - Understand branch relationships and merge history * **Staging area** - Stage specific changes or entire files for commit * **Commit templates** - Use predefined commit message templates for consistency ### **Database explorer** Automatically connects to your configured data warehouse, providing direct data access and query capabilities: IDE Database Explorer **Functionality:** * **Schema browsing** - Browse datasets, schemas, and tables in your data warehouse * **Query execution** - Execute ad-hoc SQL queries with syntax highlighting and autocomplete * **Data preview** - Preview table data and examine schema information with pagination * **Result export** - Export query results in multiple formats (CSV, JSON, Excel) * **Query management** - Save frequently used queries for reuse and organization **Database explorer benefits:** * **Real-time data access** - Query live data during development and testing * **Schema inspection** - Understand table structures and relationships * **Data validation** - Verify data quality and transformations during development ### **dbt Power User extension** Pre-installed extension providing comprehensive dbt development capabilities with seamless integration to your configured dbt environment: dbt Power User Extension **Core features:** * **Model execution** - Run individual models, selections, or entire dbt projects with integrated test runner * **Interactive lineage** - Visualize model dependencies with clickable graphs and impact analysis * **Documentation generation** - Create and view dbt docs with integrated preview and automatic refresh * **Test execution** - Run dbt tests with detailed results and error reporting * **SQL compilation** - Preview compiled SQL queries before execution **dbt development workflow:** * **One-click execution** - Run models directly from the file explorer or editor * **Dependency resolution** - Automatically handle model dependencies and execution order * **Error handling** - Clear error messages and debugging information * **Performance monitoring** - Track model execution times and resource usage ## IDE controls and session management ### **Session control options** Located in the top-left corner, IDE controls provide session management options for maintaining optimal development environment performance: IDE Session Controls ### **Restart session** **When to restart:** * After updating credentials or dbt configuration in Settings * When IDE becomes unresponsive or feels slower than usual * To refresh environment variables and connections **What happens during restart:** * **Environment refresh** - Reloads all environment variables and configurations * **Connection re-establishment** - Reconnects to data warehouse and repositories * **State preservation** - Maintains open files and unsaved changes ### **Reset IDE** ⚠️ **Use with caution:** * **Complete environment reset** - Returns IDE to default state * **Data loss warning** - Permanently deletes all custom settings and uncommitted changes * **Extension removal** - Uninstalls all custom extensions and configurations **When to consider reset:** * **Severe environment corruption** - When restart doesn't resolve issues * **Extension conflicts** - When multiple extensions cause instability * **Configuration errors** - When settings become corrupted and unrecoverable * **Clean slate development** - Starting fresh with a new project setup **Reset IDE Warning** The Reset IDE option will permanently delete all custom settings, uncommitted changes, and installed extensions. Always try restarting the session first before using the reset option. ## Terminal and command line ### **Terminal access** The IDE provides full command-line access with some important differences from standard terminals: **Important terminal shortcuts:** * **Paste into terminal**: `Cmd + Shift + V` (Mac) or `Ctrl + Shift + V` (Windows/Linux) * **Copy from terminal**: Select text and `Cmd + C` / `Ctrl + C` **Terminal capabilities:** * **Full shell access** - Use bash, zsh, or other available shells * **Package management** - Install Python packages, system tools, and development utilities * **Git operations** - Execute Git commands directly from the terminal * **Environment management** - Activate virtual environments and manage Python versions ## Customization and extensions ### **Extension marketplace** Install additional VS Code extensions to customize your development environment: **Popular extensions for data development:** * **SQL formatters** - Consistent code formatting for SQL queries * **YAML validators** - Validate dbt configuration files and YAML syntax * **Language-specific syntax highlighters** - Enhanced syntax support for various file types * **Git tools** - Enhanced version control workflows and visualizations * **Python extensions** - Improved Python development experience with IntelliSense * **Database tools** - Additional database connection and query tools **Installing extensions:** 1. Open the Extensions panel (Ctrl+Shift+X) 2. Search for the extension you want 3. Click Install to add it to your IDE 4. Restart the session if required by the extension ### **Workspace customization** **Customizable settings:** * **Themes and color schemes** - Personalize the visual appearance * **Editor preferences** - Font size, tab settings, word wrap, and more * **Keyboard shortcuts** - Customize shortcuts for common operations * **File associations** - Set default applications for different file types * **Terminal preferences** - Shell selection, font, and color schemes **Settings persistence:** * **User settings** - Apply to all your IDE sessions * **Workspace settings** - Specific to individual projects * **Extension settings** - Configure individual extension behavior ## Performance optimization ### **Maintaining IDE responsiveness** **Best practices for optimal performance:** * **Close unused files** - Keep only necessary files open to free up memory * **Limit large query results** - Use pagination in database explorer for large datasets * **Restart sessions periodically** - Refresh the environment during long development sessions ### **Memory management** **Signs of memory issues:** * **Sluggish performance** - IDE becomes slow to respond * **High memory usage** - Browser shows high memory consumption * **Extension errors** - Extensions fail to load or function properly * **Session timeouts** - Frequent session restarts or crashes **Resolution strategies:** * **Restart session** - Refresh the environment and clear memory * **Close unused tabs** - Reduce the number of open files * **Disable heavy extensions** - Temporarily disable resource-intensive extensions * **Clear browser cache** - Reset browser data if issues persist ## Next steps Now that you understand the IDE interface, you're ready to start developing: **Start coding** Learn how to build dbt models, Python applications, and Streamlit dashboards. # Migrating to VS code IDE Source: https://documentation.5x.co/core-features/ide/moving-from-legacy-ide-to-vscode A practical guide to migrate smoothly to the 5X VS Code–based IDE with setup, workflows, and best practices **Welcome to 5X IDE** The 5X IDE delivers a native VS Code experience directly within your browser, seamlessly integrating with your data warehouse and the broader 5X ecosystem. Configure access, warehouses, and environment basics. Connect BigQuery, Snowflake, Redshift, or PostgreSQL securely. Use 5X-managed or import existing Git repositories. Learn sessions, panels, and core tools. Versions, virtualenvs, and dependency management. Power User extension and terminal workflows. Performance, extensions, and common fixes. *** ## Prerequisites and setup Before using the 5X IDE, ensure you have the necessary access and permissions configured. The IDE integrates with your existing data infrastructure but requires initial configuration to establish secure connections. **Required access** * **Developer access** to your 5X platform workspace * **Data warehouse privileges** (BigQuery, Snowflake, Redshift, PostgreSQL) * **Git repository access** (if importing existing repositories) The IDE is particularly powerful for dbt development, Python data applications, and Streamlit dashboards. Understanding your primary use cases will help you configure the environment optimally from the start. *** ## Adding credentials ### Setting up data warehouse connections The 5X IDE establishes secure connections to your data warehouse using user managed service accounts. Navigate to **Settings → IDE → Credentials** Click **Add New Credential** against IDE and select your warehouse. Follow the prompts to provide authentication and dbt configuration. Credential configuration wizard in 5X IDE settings ### BigQuery configuration For BigQuery users, create a service account with appropriate permissions and upload the JSON key file. Access Google Cloud Console → IAM & Admin → Create service account (e.g., "5x-ide-dev-account"). Assign the **BigQuery Data Editor** role. Download the JSON key file. Upload the key file in **Settings → IDE → Credentials**. **Permission scope options** * **Project-level**: Provides access across all datasets (recommended for development) * **Dataset-level**: Granular access control for sensitive data environments The BigQuery Data Editor role provides the right balance of functionality and security, allowing read, query execution, table creation, and core development operations. ### Dbt development configuration During credential setup, configure your dbt development environment to ensure commands execute with correct settings. BigQuery credential form with dbt options **Configuration fields** * **dbt version**: 1.6.18, 1.7.19, 1.8.9, or 1.9.10 (use latest stable) * **project id**: Your warehouse project identifier * **dataset name**: Target dataset for development workloads (e.g., `dbt_dev_yourname`) * **target name**: Optional custom target name Once configured, the system validates your settings and establishes connections. You'll receive confirmation when setup is complete. *** ## Creating git repositories ### Repository integration with 5X platform The 5X platform maintains a registry of repositories accessible within your IDE environment, enabling automatic lineage tracking, documentation, and deployment workflows. **Key benefits** * **Automatic metadata extraction** and dependency tracking * **Integration with metrics layer** and conversational AI * **Consistent repository access** across your workspace ### Create new repositories Navigate to **Settings → IDE → Repositories → New Repository** to open the repository configuration interface. Repository creation and configuration interface ### Repository types and templates Complete dbt project template with sample models and tests. Includes a properly configured `dbt_project.yml` and starter macros. Flexible template for Python applications or mixed projects with full control over structure. Starter Streamlit app with sample visualizations for interactive data apps. ### Repository management options Creates a fully managed Git repository within your 5X workspace. * Automatic backup and integrated version control * Seamless integration with all 5X platform features * Platform-managed maintenance and access control Connect external repositories while maintaining provider relationships. * Requires SSH URL to the existing repository * Supports public and private repositories with authentication * Works with GitHub, GitLab, Bitbucket **Critical integration note** Always create repositories through **Settings → IDE → Repositories** rather than using Git commands directly in the IDE terminal. Repositories created directly within the IDE won't appear automatically in your 5X workspace registry. Properly created repositories will appear in your IDE file explorer automatically when you start new sessions. *** ## Navigating the IDE ### Starting IDE sessions The IDE provides a browser-based VS Code environment that initializes with your configured credentials, repositories, and tools. Go to the **IDE** tab in your 5X platform. Click **Start Session** and wait for some time. Environment configurations, saved changes, and extensions persist between sessions. Sessions auto-terminate after 30 minutes of inactivity. Start session button in the IDE tab ### IDE interface components IDE file explorer view **File explorer** * Navigate repository file trees and project structures * Create, edit, and delete files * Access documentation and config files * Drag-and-drop organization IDE source control panel with changes and branches **Source control** * Commit with descriptive messages * Create, switch, and merge branches * Pull and push to remotes * View diffs and history * Resolve merge conflicts visually Database explorer connected to a data warehouse **Database explorer** * Browse datasets, schemas, and tables * Execute ad-hoc SQL queries * Preview table data and schemas * Export results * Save frequently used queries dbt Power User extension in the IDE **dbt Power User extension** * Run models, selections, or entire dbt projects * Visualize lineage with interactive graphs * Generate and view dbt docs * Run dbt tests with detailed results * Preview compiled SQL IDE controls for restart and reset **IDE controls** * **Restart session** after updating credentials or dbt config, or when IDE becomes sluggish * **Reset IDE** deletes all custom settings and uncommitted changes; try restart first *** ## Working with Python ### Python environment overview The IDE comes with multiple Python versions via `pyenv` for compatibility across projects. **Available versions** * Python 3.8.20 * Python 3.9.23 * Python 3.10.18 * Python 3.11.13 (default via `PYENV_VERSION`) * Python 3.12.11 * Python 3.13.4 ```bash theme={null} ls /root/.pyenv/versions ``` ### Virtual environment management Create isolated environments per project. ```bash theme={null} # Using Python 3.11.13 (default) /root/.pyenv/versions/3.11.13/bin/python -m venv my_project_env # Using specific Python version /root/.pyenv/versions/3.10.18/bin/python -m venv legacy_project_env ``` Activate and manage environments: ```bash theme={null} # Activate environment source my_project_env/bin/activate # Verify active environment (should show your env path) which python # Deactivate when finished deactivate ``` ### Dependency management best practices Maintain dependencies with `requirements.txt` for reproducibility. ```text theme={null} # Core data processing pandas==2.0.3 numpy==1.24.3 # API and web requests requests==2.31.0 urllib3==2.0.4 # Visualization matplotlib==3.7.2 seaborn==0.12.2 # Development tools jupyter==1.0.0 pytest==7.4.0 ``` Install dependencies: ```bash theme={null} # Activate environment first source my_project_env/bin/activate # Install from requirements file pip install -r requirements.txt # Install additional packages and update requirements pip install scikit-learn==1.3.0 pip freeze > requirements.txt ``` ### Terminal usage and shortcuts The IDE terminal runs in the browser, so shortcuts differ from native terminals. * **Paste into terminal**: Cmd + Shift + V (Mac) or Ctrl + Shift + V (Windows/Linux) * **Standard Cmd/Ctrl + V** does not paste into the terminal * **Copy from terminal**: Select text and use Cmd + C / Ctrl + C Useful commands: ```bash theme={null} # Check Python version and location python --version which python # Run Python scripts python my_script.py # Start interactive Python shell python # Install packages in development mode pip install -e . ``` *** ## Working with dbt ### Dbt Power User extension (recommended) The extension automatically uses your configured dbt settings from **Settings → Credentials**, including version, connections, and targets. dbt Power User workflows in the IDE **Key workflows** * Execute individual models (`dbt run --select model_name`) or selections * Test models with the integrated test runner * Preview compiled SQL before execution * Generate and view documentation ### Command-line dbt development Terminal-based workflows are supported with pre-configured virtual environments per dbt version. ```bash theme={null} # List available dbt environments ls /root/.venv # Activate specific dbt version source /root/.venv/dbt-1.8.9/bin/activate # Verify dbt installation dbt --version ``` **Available versions** * dbt-1.6.18 (`/root/.venv/dbt-1.6.18/`) * dbt-1.7.19 (`/root/.venv/dbt-1.7.19/`) * dbt-1.8.9 (`/root/.venv/dbt-1.8.9/`) * dbt-1.9.10 (`/root/.venv/dbt-1.9.10/`) Common commands: ```bash theme={null} # Navigate to your dbt project directory cd /path/to/your/dbt/project # Run entire project dbt run # Run specific models dbt run --select staging.stg_customers+ # Test your models dbt test # Generate documentation dbt docs generate dbt docs serve ``` ### Lineage visualization Understand dependencies and data flow across your project. Interactive lineage graph for dbt models Open any model file in the editor. Open the **Lineage** tab to explore dependencies. Click nodes to move between upstream and downstream models. **Lineage features** * Interactive navigation between related models * Control dependency depth * Impact analysis for change assessment * Visual debugging to identify circular dependencies ### Dbt development best practices **Project organization** * Keep clear folders (e.g., `staging/`, `marts/`) * Consistent naming conventions * Focus model files and keep them small **Testing and validation** * Run dbt tests frequently during development * Validate transformed data in the database explorer * Preview compiled SQL to verify transformations **Documentation and collaboration** * Generate docs regularly * Commit frequently with descriptive messages * Use branches for features and manage in the source control panel *** ## Tips & troubleshooting ### Environment customization and extensions You can tailor the environment to your preferences: * Install additional VS Code extensions from the marketplace * Configure themes, color schemes, and editor preferences * Add CLI tools, aliases, and shell preferences * Customize keyboard shortcuts and workspace settings ### Performance optimization **Maintain responsiveness** * Close unused files and tabs * Limit large query result sets in the database explorer * Restart sessions periodically during long development sessions * Monitor resource usage in browser devtools if you notice slowness **Manage large projects** * Use dbt selection syntax to run subsets during development * Consider Git sparse checkout for very large repositories * Split very large monoliths into smaller, focused repos when possible ### Common issues and resolutions **IDE session won't start** * Verify credentials in **Settings → IDE → Credentials** * Ensure service account has warehouse permissions * Check repository permissions for imported repos * Contact support if initialization still fails after verification **dbt commands failing** * Activate the appropriate dbt virtual environment * Verify credential configuration matches your profile requirements * Ensure target dataset/schema exists and has write permissions * Review dbt logs for specific errors **Git operations not working** * Ensure repository was created via **Settings → IDE → Repositories** * Verify SSH key configuration for external imports * Check repository permissions (read/write) * Confirm repository URL is correct and accessible **Database connection issues** * Test credentials with the database explorer * Verify firewall rules allow connections from the 5X platform * Check service account key validity and expiry * Confirm required datasets and tables exist ### Development workflow best practices **Repository management** * Always create repositories via **Settings** for full integration * Commit frequently with meaningful messages * Use feature branches and merge back to main when complete * Manage workflows visually in the source control panel **Data development efficiency** * Build incrementally rather than full-project runs * Validate data quickly in the database explorer * Use lineage visualization to assess impact * Keep code clean, documented, and consistently formatted *** ## Conclusion As we continue to enhance the platform based on your feedback, explore the features in this guide and share your experiences. Your insights help us build the best development environment for data professionals. **Key takeaways** * Always configure credentials and repositories through **Settings** for full platform integration * Leverage the pre-installed **dbt Power User** extension for streamlined dbt development * Use virtual environments and proper dependency management for reliable Python workflows * Use integrated lineage and database exploration for better understanding and faster debugging Happy coding with 5X IDE. # Overview Source: https://documentation.5x.co/core-features/ide/overview Integrated development environment for data professionals with VS Code, dbt, Cube, Python, and Streamlit support **Focus:** Build data pipelines, models, and applications using our integrated VS Code-based IDE with pre-configured tools for dbt, Python, and Streamlit development. The 5X IDE delivers a native VS Code experience directly within your browser, seamlessly integrating with your data warehouse and the broader 5X ecosystem. Build reliable data models, create custom Python applications, and develop interactive dashboards all in one powerful development environment. ## What is the 5X IDE? The 5X IDE is a comprehensive development environment designed specifically for data professionals. It combines the power of VS Code with specialized tools for data development, providing everything you need to build, test, and deploy data applications. ### **Core capabilities** **Familiar development experience** Full VS Code functionality with extensions, themes, and customization options **Direct data connectivity** Query and explore your data warehouse with integrated database explorer **Data transformation** Build and test dbt models with lineage visualization and testing tools **Custom development** Create Python scripts, data pipelines, and machine learning applications **Interactive applications** Build and deploy interactive data dashboards and applications **Centralized data governance** Define and transform data with Metrics models **Version control** Full Git workflow with visual interface and collaboration tools ## IDE features overview ### **Pre-configured development stack** The IDE comes ready-to-use with all the tools data professionals need: **Python environments:** * Multiple Python versions (3.9.23, 3.10.18, 3.11.13, 3.12.11, 3.13.4) * Virtual environment management * Package management with pip * Interactive Python shell **dbt development:** * Pre-installed dbt Power User extension * Multiple dbt versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * Lineage visualization * Model execution and testing **Data warehouse support:** * BigQuery, Snowflake, Redshift, PostgreSQL * Secure credential management * Real-time data access * Query execution and export **Development tools:** * Git integration with visual interface * Terminal access with custom shortcuts * Extension marketplace * Customizable workspace ## Getting started ### **Quick start workflow** Configure your data warehouse credentials in Settings → IDE → Credentials Set up a new repository in Settings → IDE → Repositories Navigate to IDE tab and click "Start Session" Access your repositories and start building data applications Tip: Verify that you’ve selected a supported compute profile. ### **Supported use cases** The IDE is optimized for these common data development tasks: **Data modeling and transformation:** * Build dbt models with version control * Test data transformations * Generate documentation * Visualize data lineage **Python data applications:** * Create data processing scripts * Build machine learning pipelines * Develop API integrations * Process and analyze data **Interactive dashboards:** * Build Streamlit applications * Create data visualizations * Deploy interactive reports * Share insights with stakeholders ## Getting started Ready to start developing? Follow these quick start guides: **Configure your environment** Set up data warehouse credentials and repositories to start developing. **Start building** Learn how to create dbt models, Python applications, and Streamlit dashboards. **Switch from legacy IDE** Step-by-step guide to move to the new VS Code–based IDE. # Setting up the IDE Source: https://documentation.5x.co/core-features/ide/setup Complete guide to setting up your IDE environment with data warehouse credentials and repository configuration **Focus:** Configure your IDE environment with proper data warehouse connections, repository access, and development tools to start building data applications. Before you can start developing in the 5X IDE, you need to configure your data warehouse credentials and set up repositories. This guide walks you through the complete setup process for all supported data warehouses and repository types. **Setup Requirements** Ensure you have developer access to your 5X platform workspace and appropriate data warehouse privileges before beginning the setup process. ## Prerequisites and access requirements ### **Required platform access** Before using the 5X IDE, ensure you have the necessary access and permissions: * **Developer access** to your 5X platform workspace * **Data warehouse privileges** (BigQuery, Snowflake, Redshift, or PostgreSQL) * **Git repository access** (if importing existing repositories) The IDE is particularly powerful for dbt development, Python data applications, and Streamlit dashboards. Understanding your primary use cases will help you configure the environment optimally from the start. ### **Data warehouse requirements** Different warehouses have specific requirements: **Google Cloud setup required** Service account with BigQuery Data Editor role and JSON key file **Key pair authentication** Private key and passphrase for secure authentication **Cluster access** Cluster endpoint, database credentials, and network access **Server connection** Host, port, database name, and authentication credentials ### **Python version requirements** The IDE comes pre-installed with multiple Python versions managed through pyenv, providing flexibility for different project requirements and dependency compatibility. **Supported Python versions:** * **Python 3.8.20** - Extended legacy support for older projects * **Python 3.9.23** - Legacy support for older projects * **Python 3.10.18** - Stable version with good package compatibility * **Python 3.11.13** - Default version (set by PYENV\_VERSION) * **Python 3.12.11** - Latest stable with performance improvements * **Python 3.13.4** - Cutting-edge features and optimizations These versions are automatically available in your IDE environment. You can switch between versions when creating virtual environments or running Python applications. ### **dbt version requirements** The IDE supports multiple dbt versions to accommodate different project requirements and compatibility needs. **Supported dbt versions:** * **dbt 1.6.18** - Legacy support for older projects * **dbt 1.7.19** - Stable version with broad compatibility * **dbt 1.8.9** - Current stable version * **dbt 1.9.10** - Latest features and improvements These versions are available when configuring data warehouse credentials. Select the appropriate version based on your dbt project requirements and compatibility needs. ## Credential configuration ### **Accessing credential settings** Navigate to the IDE credential configuration: Access **Settings** from the main navigation in your 5X platform Click on **IDE** in the settings menu, then select **Credentials** Click **Add New Credential** against IDE to start the configuration wizard ### **Data warehouse setup examples** The IDE supports multiple data warehouses. Choose the setup process for your specific warehouse: **Google BigQuery setup** For BigQuery users, you'll need to create a service account with appropriate permissions: BigQuery IDE Credentials Setup Access Google Cloud Console → IAM & Admin → Create service account (e.g., "5x-ide-dev-account") Assign the **BigQuery Data Editor** role for optimal functionality and security Create and download the JSON key file for the service account Upload the key file through the 5X credential interface **Configuration fields:** * **Service Account Key**: Upload your JSON key file * **Project ID**: Your BigQuery project identifier * **dbt Version**: Choose from available versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * **Dataset Name**: Target dataset for development workloads * **Target Name**: Optional custom target name for dbt profile **Permission scope options:** * **Project-level**: Access across all datasets (recommended for development) * **Dataset-level**: Granular access control for sensitive data environments **Snowflake setup** Configure Snowflake credentials using key pair authentication: Snowflake IDE Credentials Setup Enter your Snowflake account URL and default role/warehouse Generate a key pair and configure private key authentication Enable dbt and set version, database, and schema preferences Test connection and save your Snowflake credentials **Configuration fields:** * **Snowflake Account**: Your Snowflake account URL * **Default Role**: Optional default role for connections * **Default Warehouse**: Optional default warehouse for queries * **Username**: Your Snowflake username * **Private Key**: Upload or paste your private key * **Private Key Passphrase**: Optional passphrase if key is encrypted * **dbt Version**: Choose from available versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * **Database**: Target database for dbt models * **Schema**: Target schema for dbt models **Amazon Redshift setup** Configure Redshift connection with username/password authentication: Redshift IDE Credentials Setup Enter your Redshift cluster host, port, and database information Provide username and password for database access Enable dbt and configure version, schema, and target settings Validate connection and save your Redshift credentials **Configuration fields:** * **Host**: Your Redshift cluster endpoint * **Port**: Database port (default 5439) * **Database**: Target database name * **Username**: Database username * **Password**: Database password * **dbt Version**: Choose from available versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * **Schema**: Target schema for dbt models * **Target**: Optional target environment name **PostgreSQL setup** Configure PostgreSQL connection for local or cloud instances: PostgreSQL IDE Credentials Setup Enter your PostgreSQL host, port, and database information Provide username and password for database access Enable dbt and configure version, schema, and target settings Validate connection and save your PostgreSQL credentials **Configuration fields:** * **Host**: Your PostgreSQL server endpoint * **Port**: Database port (default 5432) * **Database**: Target database name * **Username**: Database username * **Password**: Database password * **dbt Version**: Choose from available versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * **Schema**: Target schema for dbt models * **Target**: Optional target environment name ## Repository management ### **Repository integration benefits** The 5X platform maintains a registry of repositories accessible within your IDE environment. This integration enables advanced features like automatic lineage tracking, integrated documentation, and seamless deployment workflows. **Key benefits of integration:** * **Automatic metadata extraction** and dependency tracking * **Consistent repository access** across your workspace * **Seamless deployment workflows** to production environments ### **Creating new repositories** Navigate to **Settings → IDE → Repositories → New Repository** to access the repository configuration interface: IDE Repository Configuration ### **Repository types and templates** Choose the right template for your development needs: **Data transformation projects** Pre-configured dbt project structure with sample models, tests, and documentation templates. Includes properly configured dbt\_project.yml and basic macros. **Interactive applications** Basic Streamlit template with sample visualizations for data apps. Ready-to-use dashboard components and data visualization examples. **Metrics projects** Pre-configured Cube project structure with sample data models, schema files, and caching setup. Defines measures, dimensions, and pre-aggregations for building a consistent Metrics project. **Custom development** Maximum flexibility for Python applications or mixed-technology projects. Complete control over project organization and structure. ### **Repository management options** **Fully managed Git repository** Creates a fully managed Git repository within your 5X workspace: * **Automatic backup** and integrated version control * **Seamless integration** with all 5X platform features * **Platform handles** repository maintenance and access control * **No external Git provider** required **Best for:** New projects, teams wanting integrated workflows, simplified repository management **Connect external repositories** Connects external repositories while maintaining external Git provider relationships: * **Requires SSH URL** to existing repository * **Supports public and private** repositories with proper authentication * **Maintains connection** to GitHub, GitLab, Bitbucket, etc. * **Full Git workflow** preserved with external provider **Best for:** Existing projects, teams using external Git providers, maintaining current workflows ### **Repository setup process** Select from dbt, cube, general purpose, or Streamlit repository templates based on your project needs Choose between 5X managed repository or import existing repository For existing repositories, provide SSH URL and configure authentication Follow the configuration wizard to create your repository Confirm repository appears in your IDE file explorer **Critical integration note** Always create repositories through **Settings → IDE → Repositories** rather than using Git commands directly in the IDE terminal. Repositories created directly within the IDE won't appear automatically in your 5X workspace registry. ## Environment validation ### **Testing your setup** After configuring credentials and repositories, validate your setup: Use the database explorer to verify you can connect to your warehouse Confirm repositories appear in the IDE file explorer Run a simple dbt command to verify dbt is properly configured Verify Python versions and virtual environment creation works ### **Common setup issues** **Troubleshooting credential issues:** * **Verify service account permissions** for BigQuery users * **Check private key format** for Snowflake users * **Confirm network access** for Redshift/PostgreSQL * **Validate authentication credentials** for all warehouse types **Repository integration issues:** * **Ensure repository created through Settings** interface * **Check SSH key configuration** for external repositories * **Verify repository permissions** and access rights * **Confirm repository URL** is correct and accessible **dbt configuration problems:** * **Verify credential configuration** matches dbt profile requirements * **Check target dataset/schema** exists and has write permissions * **Confirm dbt version** is compatible with your warehouse * **Review dbt logs** for specific error messages ## Next steps Once your IDE is properly configured, you're ready to start developing: **Learn the interface** Navigate the IDE interface and understand core components for efficient development. **Start developing** Learn how to build dbt models, Python applications, and Streamlit dashboards. # Overview Source: https://documentation.5x.co/core-features/ingestion/Overview Fully managed data connectivity with 600+ connectors - bring all your data into a centralized warehouse in minutes **Focus:** Understand how 5X Ingestion makes it easy to unify fragmented data across your organization with fully managed, no-code data pipelines. ## What is 5X Ingestion? 5X Ingestion is a fully managed capability for bringing data from 600+ sources into a centralized warehouse. It's the foundation for activating analytics, AI, and operational use cases on the 5X platform. 5X Ingestion makes it easy to unify fragmented data across your organization. With hundreds of prebuilt connectors and a managed setup, you can move data from disparate systems into your warehouse in minutes - no pipelines to build or infrastructure to maintain. **Fully Managed Experience** 5X leverages best-in-class ingestion engines under the hood. No need to manage multiple ingestion tools or accounts. 5X Ingestion Dashboard ## Key capabilities ### **600+ out-of-the-box connectors** Connect to your entire data ecosystem with pre-built, production-ready connectors for: * **SaaS applications:** Salesforce, HubSpot, NetSuite, Zendesk, and hundreds more * **Databases:** MySQL, PostgreSQL, Oracle, SQL Server, MongoDB, and other data stores * **Cloud storage:** AWS S3, Google Cloud Storage, Azure Blob, and file systems * **APIs and webhooks:** Custom integrations for unique data sources ### **Fully managed setup** * **No third-party accounts** required - everything managed within 5X * **Automatic updates** to connectors and ingestion engines * **Infrastructure management** handled completely by 5X * **Security and compliance** built-in with enterprise-grade controls ### **Advanced features** * **Incremental and full syncs** to optimize data transfer * **Schema and table-level control** over what data to ingest * **Secure column hashing** and exclusion for sensitive data * **Built-in monitoring and alerts** for pipeline health ## Getting started Ready to connect your first data source? Here's how: **15-minute quickstart** Follow our step-by-step guide to connect your first data source and see data flowing into your warehouse. **Detailed setup process** Learn the complete connector setup flow with authentication, configuration, and best practices. ## Core documentation **Manage your connections** Monitor, troubleshoot, and optimize your data connections with comprehensive management tools. **Stay informed** Set up comprehensive monitoring and alerting for your data pipelines with real-time notifications. ## Key benefits ### **Unified data access** Break down data silos by bringing all your data into a centralized warehouse where it can be easily accessed, analyzed, and acted upon by your entire organization. ### **No infrastructure management** Focus on using your data, not managing ingestion infrastructure. 5X handles all the complexity of scaling, monitoring, and maintaining your data pipelines. ### **Rapid time to value** Get from data source to actionable insights in minutes, not weeks. Pre-built connectors and automated setup mean you can start analyzing your data immediately. ### **Enterprise-grade reliability** Built-in error handling, automatic retries, and comprehensive monitoring ensure your data flows are reliable and your business can depend on fresh, accurate data. ## Integration with 5X platform Ingestion works seamlessly with other 5X capabilities: * **Data Warehousing:** All ingested data is stored in your managed or imported warehouse * **Modeling:** dbt models can reference any ingested data source for transformations * **Orchestration(Jobs):** Workflows can trigger and coordinate ingestion jobs with other processes * **Business Intelligence:** Dashboards automatically reflect newly ingested data * **Metrics Store:** Metrics can be defined on any ingested data source * **Data & AI Apps:** Applications can access fresh data through APIs and SDKs *** # Connector Management Source: https://documentation.5x.co/core-features/ingestion/connector-management Manage your data connectors with monitoring, scheduling, and control tools for reliable data ingestion **Focus:** Learn how to monitor, control, and optimize your active data connectors through the 5X management interface. Once your connectors are set up and running, the connector management dashboard provides comprehensive tools to monitor performance, adjust scheduling, and troubleshoot issues. This guide covers all the management capabilities available in the 5X platform. ## Connector dashboard The main connector dashboard displays all your active data connections in a centralized table view: Connector Management Dashboard showing all active connections ### Status indicators Each connector displays one of four possible statuses: **Active** - Connector is running normally * Syncing data according to schedule * No issues detected **Broken** - Connector has encountered errors * Failed authentication or connection issues * Data source unavailable or changed * Requires immediate attention **Incomplete** - Connector setup is not finished * Missing required configuration * Partial setup that needs completion **Paused** - Connector is temporarily disabled * Manually paused by user * Not syncing data until resumed Filter by status to quickly identify connectors that need attention or are in specific states. ## Connector actions Access all connector management functions through the actions menu: Connector actions menu showing Pause, Sync now, Edit setup, Historical re-sync, and Delete options ### Available actions **Control sync execution** Temporarily stop or restart connector syncing without losing configuration **Trigger immediate sync** Force a sync to run immediately, outside of the regular schedule **Modify configuration** Change connector settings, authentication, or data selection **Backfill historical data** Re-sync all historical data (free operation) **Remove connector** Permanently delete the connector and stop all syncing ## Sync scheduling ### Frequency options Configure how often your connectors sync data by adjusting the sync frequency: Sync frequency selection dialog with Fixed Interval and CRON Expression options **Choose from predefined sync intervals:** * **1 minute** - For critical, rapidly changing data * **5 minutes** - High-frequency updates * **15 minutes** - Frequent refresh requirements * **1 hour** - Hourly updates * **6 hours** - Four times daily (common default) * **24 hours** - Once daily *Best for:* Most standard business use cases and regular reporting needs **Advanced scheduling with CRON expressions:** ```cron theme={null} # Every weekday at 9 AM 0 9 * * 1-5 # Every hour during business hours (9 AM - 5 PM) 0 9-17 * * * # Every 15 minutes during peak hours */15 8-18 * * * ``` *Best for:* Complex scheduling patterns and specific business requirements ## Connector details ### Individual connector view Click on any connector to access detailed management options: Individual connector detail page showing HubSpot to Snowflake connection The connector detail page shows: * **Source and destination** - Visual representation of data flow * **Status and last sync** - Current operational state * **Quick actions** - Test connection, Sync now, Pause * **Navigation tabs** - Schema and Sync History ### Test connection Use the **Test connection** button to verify connectivity and diagnose issues: Click **Test connection** to start verification System verifies credentials and permissions Tests network connectivity to source system Displays success or detailed error information Test connection button and results interface **Common test results:** * **Success** - Connection is working properly * **Authentication failed** - Credentials need updating * **Connection timeout** - Network or firewall issues * **Permission denied** - Insufficient access rights ## Sync history ### Monitoring sync performance The Sync History tab provides detailed information about all sync runs: Sync History tab showing successful sync runs with duration and timing details ### History details Each sync run displays: **Sync ID** - Unique identifier for the specific sync run **Status** - Success, Failed, or In Progress * **Success** - Sync completed without errors * **Failed** - Sync encountered errors and stopped * **In Progress** - Sync is currently running **Duration** - Total time taken for the sync (2m 28s, 41s, etc.) **Start Time** - When the sync began execution **End Time** - When the sync completed ### Performance insights Use sync history to: * **Identify trends** - Monitor sync duration over time * **Spot issues** - Find patterns in failed syncs * **Optimize timing** - Adjust frequency based on performance * **Troubleshoot** - Investigate specific sync failures Sync history is available for the last 60 days, helping you track long-term performance trends and identify optimization opportunities. ## Troubleshooting ### Common issues and solutions **Symptoms:** Test connection fails, sync status shows "Broken" **Common causes:** * Expired or changed authentication credentials * Network connectivity issues * Source system maintenance or downtime * Firewall or security policy changes **Solutions:** 1. Use **Test connection** to diagnose the specific issue 2. Update credentials in **Edit setup** if authentication failed 3. Check source system status and availability 4. Contact your IT team for network/firewall issues **Symptoms:** Increasing sync duration, frequent timeouts **Common causes:** * Growing data volume in source system * Network latency or bandwidth limitations * Source system performance degradation * Suboptimal sync frequency settings **Solutions:** 1. Review sync history for performance trends 2. Adjust sync frequency to reduce load 3. Consider filtering data to reduce volume 4. Contact support for optimization guidance **Symptoms:** Status shows "Incomplete", missing expected data **Common causes:** * Setup process not fully completed * Missing required configuration parameters * Insufficient permissions on source system * Schema changes in source system **Solutions:** 1. Use **Edit setup** to complete configuration 2. Verify all required fields are populated 3. Check source system permissions 4. Review and approve any schema changes ## Best practices ### Sync frequency optimization **Choose appropriate frequencies:** * **High-frequency syncs** for operational data that changes frequently * **Standard frequencies** for most business reporting needs * **Low-frequency syncs** for historical or slowly changing data **Consider system impact:** * Avoid overwhelming source systems with too frequent syncs * Balance data freshness needs with system performance * Use off-peak hours for large data volumes *** ## Next steps **Set up alerts** Configure notifications for sync failures and performance issues. # Connector Setup Source: https://documentation.5x.co/core-features/ingestion/connector-setup Complete guide to setting up and configuring data connectors with authentication, data selection, and optimization **Focus:** Learn the complete process of setting up data connectors from initial selection through to active data syncing. Setting up a new data connection in 5X is designed to be straightforward, but there are many configuration options and best practices to ensure optimal performance and security. This guide covers everything you need to know. ## Setup overview The connector setup process follows a consistent pattern across all data sources: ### **Setup stages** 1. **Source selection** - Choose your data source and verify compatibility 2. **Authentication** - Securely connect to your data source 3. **Data selection** - Choose which data to sync 4. **Configuration** - Set sync schedules and advanced options 5. **Testing & launch** - Verify setup and start syncing ## Source selection ### **Choose your connector** 5X provides 600+ out-of-the-box connectors presented as a searchable list: **How to find your connector:** * **Search by name** - Use the search bar to find your desired source (e.g., "Google Sheets", "Salesforce", "PostgreSQL") * **Browse the list** - Scroll through the complete list of available connectors * **Request custom connectors** - If a connector isn't available, you can request it and it's typically delivered within days **Popular connector types include:** * **SaaS applications** - Salesforce, HubSpot, NetSuite, Zendesk, Shopify, Stripe * **Databases** - PostgreSQL, MySQL, MongoDB, Redshift, Snowflake, BigQuery, Oracle * **Analytics platforms** - Google Analytics, Facebook Ads, Mixpanel, Adobe Analytics * **File storage** - Google Sheets, Amazon S3, Google Cloud Storage, FTP, SFTP If a connector is not available in the catalog, you can request custom connectors, which are typically delivered within days. You can also ingest from modern APIs or legacy sources using custom integration blueprints maintained by 5X. ### **Verify compatibility** Before proceeding, check: * **Data source version** compatibility * **Network connectivity** requirements * **Permission levels** needed * **Data volume** estimates ### **Destination selection** Each 5X workspace connects to a single data warehouse. Your destination for ingestion is configured through App Connections: **How destinations work:** * **Single warehouse per workspace** - Your workspace connects to one Snowflake account, BigQuery project(s), or Redshift cluster * **Multiple destinations within warehouse** - Create different App Connections of type "Ingestion" to ingest to different schemas, databases, or locations within your warehouse * **App Connection setup** - Destinations are created via Settings → App Connections by adding a new connection of type "Ingestion" **Supported warehouse types:** * **Snowflake** - Connect to your Snowflake account with different databases and schemas * **BigQuery** - Import multiple projects during workspace setup, then create destinations for different datasets * **PostgreSQL** - Connect to your PostgreSQL database with different schemas * **Redshift** - Connect to your Redshift cluster with different databases and schemas To set up ingestion destinations, you'll need to create App Connections of type "Ingestion" from Settings → App Connections. Refer to [Step 4: App Connections and Credentials](/quickstart/step-4-app-connections-credentials) for detailed setup instructions. ## Authentication ### **Authentication methods** 5X supports various authentication methods that differ based on your specific data source. Each connector presents the appropriate authentication options during setup: **OAuth 2.0 (Recommended when available)** * **Secure authorization** without sharing passwords * **Automatic token refresh** for uninterrupted access * **Granular permissions** control what data 5X can access * **Easy revocation** from the source system **General approach:** The OAuth flow redirects you to your source system's authorization page where you grant 5X permission to access your data. The specific fields and permissions will vary based on your data source. **Benefits:** * Enhanced security with no password sharing * Automatic credential management * Fine-grained access control * Simple credential revocation **API Keys & Tokens** * **Direct API access** using provided keys or tokens * **Static credentials** that require manual management * **Simple setup** for API-based services **General approach:** You'll need to generate credentials in your source system and provide them to 5X. The specific credential types (API keys, access tokens, bearer tokens, etc.) and required fields vary by source. **Best practices:** * Use read-only credentials when possible * Store credentials securely and rotate them regularly * Document credential permissions and expiration dates * Monitor for credential usage and anomalies **Username & Password** * **Direct database connection** using standard credentials * **Connection string** configuration for advanced setups * **SSL/TLS encryption** for secure connections **General approach:** Database connections require connection details (host, port, database name) along with user credentials. The specific connection parameters and security options vary by database type. **Security considerations:** * Use dedicated read-only database users * Restrict network access to 5X's IP ranges * Enable SSL/TLS encryption for all connections * Regularly rotate passwords and review permissions **Service Account Keys** * **Machine-to-machine** authentication * **Key files or certificates** for cloud services * **Role-based access** control **General approach:** Service accounts provide programmatic access without user interaction. The credential format (JSON keys, certificates, connection strings) depends on your cloud provider or service. **Key considerations:** * Configure minimal required permissions * Securely manage and rotate key materials * Monitor service account usage * Follow your organization's key management policies ## Data selection ### **Schema discovery** Once authenticated, 5X automatically discovers your data source structure and presents it in an organized schema view: 5X Schema Discovery Interface **What you'll see:** * **Hierarchical structure** - Your data source schema organized in a tree view * **Tables and fields** - All available tables with their individual columns/fields * **Selection controls** - Checkboxes to include or exclude specific tables and fields * **Search functionality** - Find specific tables or fields quickly * **Data indicators** - Some fields may show additional information (e.g., "Unhashed" for certain data types) ### **Table and field selection** The schema interface allows you to control exactly what data gets synced: **Selection controls:** * **Individual checkboxes** - Select or deselect specific tables and fields * **Select all tables** - Quickly enable all available tables * **Show selected tables** - Filter view to only display chosen tables * **Search tables** - Find specific tables by name ### **Field-level control** Based on your selections, 5X will sync the chosen tables and fields: **What gets synced:** * **Selected tables** - Only tables with checkboxes enabled * **Selected fields** - Only fields within those tables that are checked * **Automatic data types** - 5X handles data type mapping automatically * **Field metadata** - Some fields may include additional context (like hashing status) The exact fields and data types available depend entirely on your specific data source. 5X presents whatever structure exists in your source system, so the interface will look different for each connector type. *** ## Next steps **Manage active connectors** Learn how to monitor, troubleshoot, and optimize your existing connections. **Set up monitoring** Configure alerts and monitoring for your data pipelines. # Monitoring & Alerts Source: https://documentation.5x.co/core-features/ingestion/monitoring-alerts Set up proactive alerts for your data ingestion pipelines to stay informed about sync status, failures, and data usage **Focus:** Configure alerts to get notified about sync failures, successful completions, and data usage thresholds across your ingestion connectors. Stay on top of your data pipeline health with comprehensive alerting that notifies you through Slack channels or email when important events occur in your ingestion workflows. ## Alert setup overview 5X provides flexible alert configuration to keep you informed about your data pipeline status. You can set up alerts in two ways: **Quick connector alerts** Set up alerts directly from the ingestion dashboard for specific connectors **Centralized management** Configure and manage all alerts from a single location in Settings ## Setting up alerts ### **Method 1: From connector listing page** The quickest way to set up alerts for a specific connector is directly from the ingestion dashboard: Ingestion Dashboard with Alert Configuration Go to **Ingestion** in the main navigation to view your connectors Next to each connector, click the **icon** to open the alert configuration dialog Set up your alert preferences including: * Alert name * Connector selection (pre-selected) * Alert types (Sync Failed, Sync Successful, Row usage limit) * Delivery channels (Slack and/or Email) Click **Set alert** to activate the alert for that specific connector ### **Method 2: From settings panel** For comprehensive alert management across all connectors: Alerts Settings Panel Navigate to **Settings** from the main navigation, then click **Alerts** Click the **New Alert** button to open the alert configuration dialog Set up your alert with the following options: * **Alert name** - Descriptive name for your alert * **Type** - Select "Ingestion" for data pipeline alerts * **Connectors** - Choose which connectors to monitor * **Channel** - Select Slack channels for notifications * **Email** - Add email addresses for notifications * **Alert types** - Choose specific alert conditions Click **Set alert** to create and activate your new alert ### **Editing existing alerts** To modify an existing alert from the Settings panel: In **Settings → Alerts**, locate the alert you want to modify Click on the alert name or row to open the edit dialog Modify any alert settings including connectors, channels, or alert types Click **Update alert** to save your changes ## Alert types Configure different alert types based on your monitoring needs: **Get notified when syncs fail** Receive immediate notifications when data synchronization fails for your connectors. **When this alert triggers:** * Sync encounters an error and cannot complete * Connection to data source fails * Authentication or permission issues occur * Data transformation errors during processing **Best practices:** * Set up for all critical connectors * Route to appropriate team channels * Include escalation for repeated failures **Confirmation of successful syncs** Stay informed when data synchronization completes successfully. **When this alert triggers:** * Sync completes without errors * All data successfully loaded to warehouse * Schema changes detected and handled * Scheduled sync runs on time **Best practices:** * Use for critical business data * Consider digest format for frequent syncs * Useful for compliance and audit trails * Monitor sync completion times **Monitor data volume thresholds** Get alerts when data processing approaches or exceeds usage limits. **When this alert triggers:** * Combined row count across connectors exceeds threshold * Individual connector processes excessive data * Unexpected data volume spikes detected * Approaching monthly or daily limits **Configuration options:** * Set threshold in millions of rows (e.g., 20 million) * Daily reset monitoring * Percentage-based warnings (e.g., 80% of limit) * Connector-specific limits **Best practices:** * Set threshold below actual limits * Monitor for data anomalies * Plan for business growth * Review limits regularly ## Alert configuration best practices ### **Alert organization** Structure your alerts for maximum effectiveness: **Focused monitoring** Set up dedicated alerts for critical connectors with specific requirements **Proper escalation** Route alerts to appropriate teams based on connector ownership **Prioritization** Use different channels for different alert severities **Alert fatigue** Avoid overwhelming teams with too many notifications ### **Alert naming conventions** Use consistent naming for better organization: **Naming structure:** * **Environment prefix** - `[PROD]`, `[STAGING]`, `[DEV]` * **Connector name** - Clear identifier for the data source * **Alert type** - `Sync Failed`, `Usage Limit`, etc. * **Team identifier** - Responsible team or department **Examples:** * `[PROD] Salesforce - Sync Failed - Sales Team` * `[PROD] PostgreSQL - Usage Limit - Data Engineering` * `[STAGING] BigQuery - Sync Successful - QA Team` *** # API credentials Source: https://documentation.5x.co/core-features/metric-store/api-credentials Learn how to access SQL and REST API credentials for connecting BI tools, applications, and SQL clients to your cubes **Focus:** Understand how to access and use API credentials to connect BI tools, custom applications, and SQL clients to your Metrics Store cubes. API credentials provide secure access to your cubes through SQL and REST APIs. These credentials enable you to connect external tools, applications, and clients to query your cube data programmatically. ## Accessing API credentials ### **Opening credentials drawer** 1. **Click "API Credentials" button** in the Metrics Store header 2. **Credentials drawer opens** showing available APIs 3. **View credentials** for SQL API and REST API ### **Prerequisites** API credentials are available when: * ✅ A project is selected * ✅ A branch is selected * ✅ Server heartbeat status is **"RUNNING"** * ✅ Server has been initialized **Credentials availability** If the API Credentials button is disabled or shows no data, ensure the server heartbeat has reached **"RUNNING"** status. Credentials are generated after the server successfully starts. ## SQL API credentials ### **Overview** SQL API provides PostgreSQL-compatible connection for BI tools and SQL clients: * **PostgreSQL protocol** - Standard SQL interface * **BI tool integration** - Connect Tableau, Power BI, Looker, etc. * **SQL client support** - Use with any PostgreSQL-compatible client * **Direct querying** - Execute SQL queries against your cubes ### **Connection components** **Connection string** * Complete PostgreSQL connection URL * Includes all connection parameters * Ready to use in connection dialogs **Individual credentials:** * **Host** - Server hostname or IP address * **Port** - Database port number * **Database** - Database name * **Username** - Authentication username * **Password** - Authentication password ### **Using SQL API credentials** #### **Connect BI tools** 1. **Copy connection string** or individual credentials 2. **Open BI tool connection dialog** 3. **Select PostgreSQL connection type** 4. **Enter credentials** (host, port, database, username, password) 5. **Test connection** and save #### **Connect SQL clients** Use any PostgreSQL-compatible SQL client: * **pgAdmin** - PostgreSQL administration tool * **DBeaver** - Universal database tool * **TablePlus** - Modern database client * **Command line** - `psql` or other CLI tools Example connection: ```bash theme={null} psql -h -p -U -d ``` #### **Example SQL queries** ```sql theme={null} -- Query cube measures and dimensions SELECT orders.status, orders.total_revenue, orders.count FROM orders WHERE orders.created_at >= '2024-01-01' GROUP BY orders.status ORDER BY orders.total_revenue DESC; -- Time-based analysis SELECT orders.created_at, orders.total_revenue FROM orders WHERE orders.created_at >= '2024-01-01' AND orders.created_at < '2024-02-01' GROUP BY orders.created_at ORDER BY orders.created_at; ``` ## REST API credentials ### **Overview** REST API provides RESTful access for application integration: * **HTTP-based** - Standard REST interface * **JSON requests/responses** - Easy integration * **Token authentication** - Secure API access * **Flexible querying** - Build queries programmatically ### **Connection components** **Base URL** * API endpoint base URL * Used as prefix for all API requests * Example: `https://api.example.com/v1` **Authentication token** * Bearer token for API authentication * Include in request headers * Keep secure and rotate periodically ### **Using REST API credentials** #### **API request format** ```bash theme={null} curl -X POST 'https://your-api-url/v1/load' \ -H 'Authorization: Bearer YOUR_TOKEN' \ -H 'Content-Type: application/json' \ -d '{ "query": { "measures": ["orders.total_revenue"], "dimensions": ["orders.status"], "timeDimensions": [{ "dimension": "orders.created_at", "granularity": "day", "dateRange": ["2024-01-01", "2024-01-31"] }] } }' ``` #### **Example queries** **Simple query:** ```json theme={null} { "query": { "measures": ["orders.total_revenue"], "dimensions": ["orders.status"] } } ``` **Time-based query:** ```json theme={null} { "query": { "measures": ["orders.total_revenue"], "timeDimensions": [{ "dimension": "orders.created_at", "granularity": "month" }] } } ``` **Filtered query:** ```json theme={null} { "query": { "measures": ["orders.total_revenue"], "dimensions": ["orders.status"], "filters": [{ "dimension": "orders.status", "operator": "equals", "values": ["completed"] }] } } ``` #### **Integration examples** **Python:** ```python theme={null} import requests url = "https://your-api-url/v1/load" headers = { "Authorization": "Bearer YOUR_TOKEN", "Content-Type": "application/json" } data = { "query": { "measures": ["orders.total_revenue"], "dimensions": ["orders.status"] } } response = requests.post(url, json=data, headers=headers) results = response.json() ``` **JavaScript:** ```javascript theme={null} const url = 'https://your-api-url/v1/load'; const headers = { 'Authorization': 'Bearer YOUR_TOKEN', 'Content-Type': 'application/json' }; const data = { query: { measures: ['orders.total_revenue'], dimensions: ['orders.status'] } }; fetch(url, { method: 'POST', headers: headers, body: JSON.stringify(data) }) .then(response => response.json()) .then(results => console.log(results)); ``` ## Server status ### **Status indicator** The credentials drawer shows server status: * **🟢 Green** - Server is running and operational * **🔴 Red** - Server is stopped or unavailable * **🟡 Yellow** - Server is starting up (PENDING) ### **Server restart** If schema changes occur, you may need to restart the API server: 1. **Click "Restart" button** in credentials drawer 2. **Wait for restart** - Server will restart and reload schema 3. **Status updates** - Green indicator when ready 4. **Credentials refresh** - May need to refresh credentials after restart **When to restart** Restart the API server when: * Schema files are updated and deployed * Cube definitions are modified * Connection issues occur * Server status shows errors ## Security best practices ### **Credential management** **Protect credentials** Never share API credentials publicly. Store them securely and use environment variables or secret management tools. **Periodic rotation** Rotate API credentials periodically to maintain security. Update connected applications when credentials change. **Limit access** Only share credentials with authorized users and applications. Use role-based access control where possible. **Track access** Monitor API usage and access patterns to detect unusual activity or potential security issues. ### **Connection security** * **Use HTTPS** - Always use encrypted connections * **Token expiration** - Configure token expiration policies * **IP whitelisting** - Restrict access to known IP addresses * **Rate limiting** - Implement rate limits to prevent abuse * **Audit logging** - Track API access and usage ## Troubleshooting ### **Common issues** **Possible causes:** * Server heartbeat not RUNNING * Project or branch not selected * Server not initialized **Solutions:** * Wait for heartbeat to reach RUNNING status * Verify project and branch are selected * Check server status indicator * Try refreshing the page **Possible causes:** * Incorrect credentials * Network connectivity issues * Server not running * Firewall blocking connections **Solutions:** * Verify credentials are correct (copy again) * Check network connection * Verify server status is RUNNING * Check firewall and security group settings * Test connection from different network **Possible causes:** * Invalid or expired token * Incorrect authorization header * Token format issues **Solutions:** * Verify token is correct and not expired * Check authorization header format: `Bearer YOUR_TOKEN` * Refresh credentials if token appears invalid * Contact support if issues persist ### **Debugging connections** **Test SQL connection:** ```bash theme={null} # Test PostgreSQL connection psql -h -p -U -d # Or test with connection string psql "postgresql://username:password@host:port/database" ``` **Test REST API:** ```bash theme={null} # Test API endpoint curl -X POST 'https://your-api-url/v1/load' \ -H 'Authorization: Bearer YOUR_TOKEN' \ -H 'Content-Type: application/json' \ -d '{"query": {"measures": ["orders.count"]}}' ``` *** **Sync BI tools** Learn how to set up automatic synchronization with BI tools. **Build queries** Use Cube Playground to build and test queries before connecting external tools. # BI integration Source: https://documentation.5x.co/core-features/metric-store/bi-integration Set up automatic synchronization of cube definitions with your BI tools for consistent metrics across all analytics platforms **Focus:** Learn how to configure BI integration to automatically sync your Cube definitions with BI tools, ensuring consistent metrics across Tableau, Power BI, Looker, and other analytics platforms. BI Integration allows you to automatically synchronize your Metrics Store cube definitions with your BI tools, ensuring that all your analytics platforms use the same metric definitions and business logic. This eliminates discrepancies and ensures consistency across your entire analytics stack. ## Overview ### **What is BI integration?** BI Integration automatically syncs your Cube definitions (cubes, measures, dimensions) with connected BI tools, ensuring: * **Consistent metrics** - Same definitions across all tools * **Automatic updates** - Changes propagate automatically * **Reduced manual work** - No need to manually update each tool * **Single source of truth** - Metrics Store as the authoritative source ### **Supported BI tools** * **Metabase** - Connect open-source BI platform More tools coming soon. ## Accessing BI integration ### **Open BI integration drawer** 1. **Click "BI Integration" button** in Metrics Store header 2. **BI Integration drawer opens** showing: * List of connected BI tools * Sync status for each connection * Last sync time and status * Management actions ### **Prerequisites** Before setting up BI integration: * ✅ Active project selected * ✅ Branch selected * ✅ Server heartbeat status is **"RUNNING"** * ✅ Cube definitions are deployed ## Creating BI connections ### **Add new connection** 1. **Click "New BI Connection"** in BI Integration drawer 2. **Select BI tool type** from available options 3. **Configure connection details:** * Connection name * Tool-specific settings * Authentication credentials * Connection parameters 4. **Save connection** * Connection is created and saved * Initial sync is triggered automatically * Connection appears in connections list ### **Connection configuration** **Connection name** * Descriptive name for identification * Example: "Tableau Production", "Power BI Analytics" * Helps identify connection purpose **BI tool selection** * Choose from supported BI tools * Each tool has specific configuration requirements * Tool-specific connectors handle synchronization **Authentication** * Provide credentials for BI tool access * May include API keys, tokens, or OAuth * Stored securely and encrypted **Connection settings** * Tool-specific configuration options * Sync frequency settings * Data source mappings * Additional parameters as needed ## Managing connections ### **Connection list** The BI Integration drawer displays: * **Connection name** - Identifies each connection * **BI tool type** - Shows which tool is connected * **Sync status** - Current sync status indicator * **Last sync time** - When last sync occurred * **Actions** - Sync, edit, delete options ### **Sync status indicators** **SUCCESS** 🟢 * Last sync completed successfully * Connection is active and working * Data is up-to-date **FAILED** 🔴 * Last sync encountered errors * Connection may need attention * Review error messages for details **PENDING** 🟡 * Sync is currently in progress * Wait for completion * Status will update when done **NEVER\_SYNCED** ⚪ * Connection has never been synced * Initial sync may be needed * Set up connection and sync ### **Connection actions** **Sync now** * Manually trigger sync for specific connection * Useful after schema changes * Updates BI tool with latest definitions **Sync all** * Sync all connections at once * Convenient for bulk updates * Useful after major schema changes **Edit connection** * Update connection settings * Modify authentication credentials * Change configuration parameters **Delete connection** * Remove connection from Metrics Store * Requires confirmation * Cannot be undone ## Sync operations ### **Individual sync** Sync a specific BI connection: 1. **Locate connection** in connections list 2. **Click "Sync now"** button for that connection 3. **Wait for sync to complete** 4. **Status updates** to show result (SUCCESS or FAILED) **When to use:** * After making changes to specific cubes * Testing connection configuration * Troubleshooting connection issues * Updating single BI tool ### **Sync all** Sync all BI connections at once: 1. **Click "Sync all" button** in drawer header 2. **All connections sync** simultaneously 3. **Status updates** for each connection 4. **Results displayed** when complete **When to use:** * After major schema changes * After deploying new cube definitions * Ensuring all tools are up-to-date * Regular maintenance syncs ## Sync workflow ### **What gets synced** When you sync a BI connection, the following are synchronized: **Cube definitions** * Cube schemas and structures * Measure definitions and calculations * Dimension definitions and types * Time dimension configurations **Metadata** * Cube descriptions and documentation * Measure labels and descriptions * Dimension hierarchies * Business logic definitions **Updates** * New cubes are added * Modified cubes are updated * Deleted cubes are removed * Changes propagate automatically ### **Sync process** 1. **Initiate sync** - User triggers sync (manual or automatic) 2. **Connect to BI tool** - Establish connection using credentials 3. **Compare definitions** - Identify differences between Metrics Store and BI tool 4. **Apply changes** - Update BI tool with latest definitions 5. **Verify sync** - Confirm changes were applied successfully 6. **Update status** - Record sync status and timestamp ## Best practices ### **Connection management** **Clear identification** Use descriptive connection names that indicate purpose, environment, or team. **Stay updated** Sync regularly, especially after schema changes, to keep all BI tools current. **Track sync health** Monitor sync status regularly and address failures promptly to maintain consistency. **Validate setup** Test connections after setup and after major changes to ensure everything works correctly. ### **Sync strategy** **After schema changes:** * Sync immediately after deploying cube changes * Test in one BI tool first, then sync others * Verify sync status before considering complete **Regular maintenance:** * Schedule regular syncs (daily or weekly) * Review sync status during maintenance windows * Address any failed syncs promptly **Environment separation:** * Use separate connections for dev, staging, and production * Sync dev environment more frequently for testing * Be cautious with production syncs ### **Error handling** * **Review error messages** - Understand why syncs fail * **Check credentials** - Verify authentication is still valid * **Test connectivity** - Ensure BI tool is accessible * **Verify schema** - Check that cube definitions are valid * **Contact support** - If issues persist after troubleshooting ## Troubleshooting ### **Common issues** **Possible causes:** * Invalid or expired credentials * BI tool connectivity issues * Schema validation errors * Permission problems **Solutions:** * Verify credentials are correct and not expired * Check BI tool is accessible and running * Review schema for errors or inconsistencies * Verify connection has necessary permissions * Check error messages for specific issues **Possible causes:** * Large number of cubes to sync * Network latency issues * BI tool performance problems * Complex schema definitions **Solutions:** * Be patient for large syncs * Check network connectivity * Verify BI tool performance * Consider syncing in smaller batches **Possible causes:** * Connection not configured correctly * Authentication failures * BI tool not accessible * Incorrect connection parameters **Solutions:** * Review connection configuration * Verify credentials and authentication * Test BI tool accessibility * Check connection parameters match requirements * Try recreating connection ### **Sync status investigation** **Check sync details:** * Review last sync time * Check sync status indicator * Look for error messages * Verify connection is active **Test connection:** * Try manual sync * Verify BI tool is accessible * Check credentials * Review connection logs **Verify schema:** * Ensure cubes are properly defined * Check for validation errors * Verify schema is deployed * Test queries in Cube Playground *** **Connect applications** Learn how to use API credentials to connect custom applications. **Manage projects** Understand project and branch management for your cubes. # Cube Playground Source: https://documentation.5x.co/core-features/metric-store/cube-playground Master the Cube Playground interface for building queries, defining cube schemas, and exploring your data **Focus:** Learn how to use the Cube Playground to visually build queries, write SQL, define cube schemas, and explore your data interactively. The Cube Playground is an interactive development environment embedded in the Metrics Store that provides a visual interface for building queries, defining cube schemas, and exploring your data. It's powered by Cube.js and integrated seamlessly with the 5X platform. ## Accessing Cube Playground ### **Prerequisites** Before accessing Cube Playground: 1. **Select a project** - Choose an active Cube project 2. **Select a branch** - Choose a branch to work with ## Cube Playground interface ### **Main components** The Cube Playground interface includes: **Query builder** * Visual drag-and-drop interface for building queries * Select measures and dimensions from available cubes * Apply filters and configure aggregations * Preview query results **SQL editor** * Direct SQL query execution * Syntax highlighting and autocomplete * Query history and saved queries * Execute and test queries **Schema explorer** * Browse available cubes, measures, and dimensions * View cube definitions and metadata * Understand data structure and relationships * Navigate hierarchical cube organization **Query results** * Display query results in tables or charts * Export results to various formats * Analyze data patterns and insights **Schema files** * Edit cube definitions using YAML or JavaScript * Version-controlled schema files * Deploy and test schema changes ## Building queries ### **Visual query builder** Use the visual query builder to create queries without writing SQL: 1. **Select cube** * Choose a cube from the sidebar * View available measures and dimensions 2. **Choose measures** * Select quantitative metrics to calculate * Examples: revenue, orders, users, conversions * Can select multiple measures 3. **Add dimensions** * Select attributes for grouping and analysis * Examples: date, region, product category * Can add multiple dimensions 4. **Apply filters** * Filter data by specific values or ranges * Date range filters * Categorical filters * Numeric filters 5. **Configure aggregations** * Set time granularity (day, week, month, etc.) * Configure sorting and ordering * Set result limits 6. **Execute query** * Click **"Run"** to execute the query * View results in the results panel * Export or save query if needed ### **SQL editor** Write SQL queries directly for advanced use cases: 1. **Open SQL editor** * Switch to SQL editor tab * Write SQL queries using Cube.js SQL syntax 2. **Query syntax** ```sql theme={null} SELECT orders.status, orders.total_revenue FROM orders WHERE orders.created_at >= '2024-01-01' GROUP BY orders.status ORDER BY orders.total_revenue DESC ``` 3. **Execute and test** * Click **"Run"** to execute * Review results and query performance * Iterate and refine queries 4. **Save queries** * Save frequently used queries * Share queries with team members * Build query library ## Defining cubes ### **Cube schema structure** Cubes are defined using YAML or JavaScript schema files: ```yaml theme={null} cubes: - name: orders sql_table: orders measures: - name: count type: count - name: total_revenue sql: sum(amount) type: sum - name: average_order_value sql: sum(amount) / count(*) type: number dimensions: - name: status sql: status type: string - name: created_at sql: created_at type: time - name: customer_id sql: customer_id type: string primary_key: true ``` ### **Schema components** **Measures** * Quantitative metrics and calculations * Types: count, sum, avg, min, max, number * Custom SQL expressions * Aggregation functions **Dimensions** * Attributes for analysis and grouping * Types: string, number, time, boolean * Primary keys for joins * Hierarchical relationships **Time dimensions** * Special dimensions for time-based analysis * Support for time grains (day, week, month, etc.) * Built-in time intelligence functions * Period-over-period comparisons ### **Editing schemas** 1. **Access schema files** * Navigate to schema files in Cube Playground * Edit YAML or JavaScript files directly 2. **Make changes** * Add new cubes, measures, or dimensions * Modify existing definitions * Update SQL expressions 3. **Save and deploy** * Save schema changes * Deploy to make changes available * Changes take effect immediately 4. **Test changes** * Test queries in Playground * Verify results are correct * Iterate and refine ## Exploring data ### **Schema explorer** Browse your cube structure: 1. **View cubes** * See all available cubes in sidebar * Navigate cube hierarchy * Understand cube relationships 2. **Explore measures** * View available measures per cube * Understand measure definitions * See measure types and calculations 3. **Explore dimensions** * View available dimensions per cube * Understand dimension types * See dimension relationships ### **Data preview** Preview actual data: 1. **Sample data** * View sample data from cubes * Understand data structure * Verify data quality 2. **Query results** * Execute queries and view results * Analyze data patterns * Export results for further analysis ## Best practices ### **Query building** **Build incrementally** Start with basic queries and add complexity gradually. Test each step before adding more measures or dimensions. **Limit data scope** Apply filters to reduce data volume and improve query performance. Filter early in your query building process. **Validate before deploying** Always test queries in Cube Playground before using them in production dashboards or applications. **Add comments** Document complex queries with comments explaining business logic and calculations. ### **Schema development** * **Version control** - Commit schema changes regularly with descriptive messages * **Incremental changes** - Make small, testable changes rather than large refactors * **Documentation** - Add descriptions and examples to cube definitions * **Testing** - Test schema changes thoroughly before deploying * **Review** - Use pull requests for schema changes in protected branches ### **Performance optimization** * **Efficient measures** - Use appropriate aggregation types * **Indexed dimensions** - Ensure frequently used dimensions are indexed * **Query optimization** - Review query performance and optimize slow queries * **Data filtering** - Apply filters to reduce data volume * **Caching** - Leverage Cube.js caching for frequently accessed data ## Troubleshooting ### **Common issues** **Possible causes:** * Heartbeat status not RUNNING * Project not selected * Branch not selected * Network connectivity issues **Solutions:** * Check heartbeat status indicator * Verify project and branch are selected * Wait for server to start (check deployment options) * Refresh the page * Check browser console for errors **Possible causes:** * Invalid cube or measure names * SQL syntax errors * Missing dimensions or filters * Data source connection issues **Solutions:** * Verify cube and measure names are correct * Check SQL syntax * Ensure required dimensions are included * Verify data source connection * Review error messages for specific issues **Possible causes:** * Large data volumes * Inefficient queries * Missing indexes * Complex calculations **Solutions:** * Apply filters to reduce data volume * Optimize query structure * Check database indexes * Simplify complex calculations * Use pre-aggregations where possible *** **Connect tools** Learn how to use API credentials to connect BI tools and SQL clients. **Manage projects** Understand how to manage projects and branches for your cubes. # Deployment options Source: https://documentation.5x.co/core-features/metric-store/deployment-options Configure when your Cube service runs with always-on, on-demand, or scheduled deployment modes **Focus:** Understand deployment options for your Cube service, including always-on, on-demand, and scheduled modes, and how to configure them for your specific use case. Deployment options control when your Cube service runs and how it manages compute resources. Choose the right deployment mode based on your requirements for availability, cost, and use case. ## Deployment modes overview ### **Available options** Metrics Store supports multiple deployment modes: * **Always on** - Service runs continuously * **On demand** - Service starts when needed * **Scheduled** - Service runs on a time-based/CRON-based schedule Each mode has different characteristics and use cases. ## Always on deployment ### **Overview** **Always on** deployment keeps your Cube service running continuously, providing constant availability for queries and API access. **Characteristics:** * ✅ Service runs 24/7 * ✅ Immediate availability for queries * ✅ No startup delays * ✅ Consistent performance * ⚠️ Higher compute costs * ⚠️ Resources always allocated ### **When to use** **Production environments:** * Critical business applications * High-traffic dashboards * Real-time analytics requirements * SLA-driven use cases **Always-available analytics:** * Executive dashboards * Customer-facing analytics * Operational reporting * Time-sensitive queries ### **Configuration** 1. **Access project settings** * Click **"Settings"** button in header * Or click project name → **"Edit project"** 2. **Select deployment mode** * Choose **"Always on"** from deployment options * Save changes 3. **Service starts** * Service starts immediately * Heartbeat status reaches **"RUNNING"** * Available for queries right away **Production recommendation** Always on deployment is recommended for production environments where consistent availability and performance are critical. ## On demand deployment ### **Overview** **On demand** deployment starts the Cube service when it's needed and stops it when idle, optimizing resource usage and costs. **Characteristics:** * ✅ Cost-effective for low usage * ✅ Automatic start on first request * ✅ Automatic stop when idle * ⚠️ Startup delay on first request * ⚠️ Cold start time required ### **When to use** **Development and testing:** * Development environments * Testing and QA * Experimentation * Non-critical workloads **Low-traffic scenarios:** * Occasional analytics needs * Ad-hoc reporting * Personal dashboards * Cost-sensitive use cases ### **Configuration** 1. **Access project settings** * Go to project settings * Navigate to deployment options 2. **Select deployment mode** * Choose **"On demand"** from options * Save changes 3. **Service behavior** * Service starts when first query arrives * Runs while active * Stops automatically after idle period * Restarts on next request **Startup time** On demand deployment requires a startup time when the service starts. Users may experience a delay on the first request after an idle period. ## Scheduled deployment ### **Overview** **Scheduled** deployment runs your Cube service on a specific schedule, such as business hours or specific days of the week. **Characteristics:** * ✅ Runs during specified times * ✅ Stops automatically outside schedule * ✅ Cost optimization for time-based needs * ✅ Predictable availability windows * ⚠️ Not available outside schedule * ⚠️ Requires schedule configuration ### **When to use** **Business hours operations:** * Office hours analytics * Business day reporting * Time-zone specific needs * Cost-optimized production **Scheduled reporting:** * Daily reports * Weekly analytics * Monthly dashboards * Batch processing windows ### **Configuration** 1. **Access project settings** * Go to project settings * Select **"Scheduled"** deployment mode 2. **Configure schedule** * Choose schedule type: * **Whole Week** - Monday-Sunday with time range * **Only Weekdays** - Monday-Friday with time range * **CRON Expression** - Cron-based scheduling 3. **Set time range** * **Start time** - When service should start (e.g., 09:00) * **End time** - When service should stop (e.g., 18:00) * **Timezone** - Auto-detected or manually selected 4. **Save configuration** * Schedule is saved and activated * Service starts/stops according to schedule ### **Schedule types** **Monday-Sunday schedule:** * Runs all week with time range * Configure start and end times * Example: 08:00 - 20:00 Monday-Sunday * Stops outside time range **Use cases:** * Extended hours operations * Daily time windows * Regular availability periods **Monday-Friday schedule:** * Runs Monday through Friday * Configure start and end times * Example: 09:00 - 18:00 Monday-Friday * Stops automatically on weekends **Use cases:** * Business hours operations * Office hours analytics * Workday reporting **Cron-based scheduling:** * Advanced scheduling with cron expressions * Full flexibility for complex schedules * Start and end cron expressions * Example: `0 9 * * 1-5` (9 AM weekdays) **Use cases:** * Complex scheduling needs * Multiple time windows * Irregular schedules * Advanced automation ### **Timezone configuration** **Auto-detection:** * System automatically detects browser timezone * Uses detected timezone for schedule * Convenient default for most users **Manual selection:** * Override auto-detected timezone * Select from available timezones * Useful for remote teams or specific regions **Timezone format:** * IANA timezone format (e.g., `America/New_York`) * Supports all major timezones * Handles daylight saving time automatically ## Comparing deployment modes ### **Quick comparison** | Feature | Always On | On Demand | Scheduled | | ---------------- | ---------- | ------------- | ---------------------------- | | **Availability** | 24/7 | When needed | On schedule | | **Startup time** | None | 30-60 seconds | None (within schedule) | | **Cost** | Highest | Lowest | Medium | | **Use case** | Production | Development | Business hours | | **Performance** | Consistent | Cold starts | Consistent (within schedule) | ### **Choosing the right mode** **Select Always On if:** * Production environment with high availability needs * Real-time analytics requirements * Consistent performance is critical * Cost is not a primary concern **Select On Demand if:** * Development or testing environment * Low-traffic or occasional use * Cost optimization is important * Startup delays are acceptable **Select Scheduled if:** * Business hours operations * Predictable usage patterns * Cost optimization for time-based needs * Availability during specific windows is sufficient ## Best practices ### **Deployment strategy** **Different modes per environment** Use Always On for production, On Demand for development, and Scheduled for staging. **Right-size deployment** Choose deployment mode based on actual usage patterns, not just availability needs. **Track patterns** Monitor usage patterns to optimize deployment mode and schedule configuration. **Scale appropriately** Start with On Demand or Scheduled, upgrade to Always On as usage grows. ### **Configuration tips** * **Test deployment changes** - Verify service starts/stops correctly * **Monitor heartbeat status** - Ensure service is running when expected * **Set appropriate schedules** - Align with actual usage patterns * **Consider timezones** - Configure schedules for your team's timezone * **Review periodically** - Adjust deployment mode as needs change ## Troubleshooting ### **Common issues** **Possible causes:** * Deployment mode configuration issues * Schedule not configured correctly * Environment or resource constraints * Service startup errors **Solutions:** * Verify deployment mode is set correctly * Check schedule configuration (if scheduled) * Review environment configuration * Check server logs for errors * Try changing deployment mode **Possible causes:** * Schedule configuration (if scheduled) * Idle timeout (if on demand) * Resource constraints * Service errors **Solutions:** * Check schedule configuration and timezone * Verify on demand idle timeout settings * Review resource limits * Check service logs for errors * Consider switching to Always On **Possible causes:** * Incorrect schedule configuration * Timezone mismatch * Cron expression errors * Service not respecting schedule **Solutions:** * Verify schedule type and time range * Check timezone configuration * Validate cron expressions (if custom) * Test schedule configuration * Review schedule in project settings *** **Manage projects** Learn how to configure deployment options in project settings. **Quick start** Understand deployment options when creating your first project. # Getting started Source: https://documentation.5x.co/core-features/metric-store/getting-started Quick start guide to creating your first Cube project and accessing the Cube Playground **Focus:** Get up and running with Metrics Store by creating your first project, accessing Cube Playground, and querying your first cubes. This guide will walk you through creating your first Cube project and getting started with the Metrics Store in just a few minutes. ## Prerequisites Before you begin, ensure you have: * Access to the 5X Platform * Metrics Store permissions * An environment configured in your workspace ## Step 1: Navigate to Metrics Store 1. **Go to Metrics Store page** * Navigate to `/metrics-store` in your browser * Or click **"Metrics Store"** from the navigation menu 2. **View the interface** * If you have no projects, you'll see a **"Create project"** button * If projects exist, use the project dropdown in the header and select **"New project"** ## Step 2: Create your first project ### **Project details** Fill in the project creation form: **Project name** * Choose a unique, descriptive name (max 50 characters) * Example: `My First Cube Project` **Repository type** * **Platform managed** (recommended for beginners) - 5X-managed Git repository * **Imported** - Connect an existing GitHub repository **Environment** * Select your environment from the dropdown * Choose the environment where your data warehouse is configured **Deployment option** * **Always on** - Service runs continuously (recommended for production) * **On demand** - Service starts when needed * **Scheduled** - Time-based or CRON-based scheduling ### **Add project** Click **"Add Project"** and wait for project creation: * Repository cloning will begin automatically (if imported) * Project initialization takes a few seconds * You'll be automatically taken to the project workspace **Platform managed repositories** For beginners, we recommend starting with **Platform Managed** repositories. 5X handles all Git repository management, making setup easier and faster. ## Step 3: Access Cube Playground Once your project is created: 1. **Cube Playground loads automatically** in the main workspace area 2. **Start exploring** your cubes and data The Cube Playground is an interactive interface where you can: * Build queries visually using the query builder * Write SQL queries directly * Explore available cubes, measures, and dimensions * Test queries before deploying ## Step 4: Get API credentials To connect BI tools or applications to your cubes: 1. **Click "API Credentials"** button in the header 2. **View SQL API credentials**: * Copy the connection string * Use in BI tools or SQL clients * Credentials include host, port, database, username, and password 3. **Switch to REST API tab** for REST API credentials: * Base URL and authentication token * Use for RESTful API access **Credentials availability** API credentials are only available when: * A project is selected * A branch is selected * Server heartbeat status is **"RUNNING"** ## Step 5: Create your first cube In the Cube Playground, you can define cubes using YAML or JavaScript. Here's a simple example: ```yaml theme={null} cubes: - name: orders sql_table: orders measures: - name: count type: count - name: total_revenue sql: sum(amount) type: sum dimensions: - name: status sql: status type: string - name: created_at sql: created_at type: time ``` **Key components:** * **Measures** - Quantitative metrics (counts, sums, averages) * **Dimensions** - Attributes for analysis (status, date, category) * **Time dimensions** - Special dimensions for time-based analysis Save your cube definition and deploy it to make it available for querying. ## Step 6: Query your data ### **Using SQL API** Connect a SQL client or BI tool using the SQL API credentials: ```sql theme={null} SELECT orders.status, orders.total_revenue FROM orders GROUP BY orders.status ``` ### **Using REST API** Use the REST API credentials to query data programmatically: ```bash theme={null} curl -X POST 'https://your-api-url/v1/load' \ -H 'Authorization: Bearer YOUR_TOKEN' \ -d '{ "query": { "measures": ["orders.total_revenue"], "dimensions": ["orders.status"] } }' ``` ### **Using Cube Playground** Build queries visually in the Cube Playground: 1. Select your cube from the sidebar 2. Choose measures and dimensions 3. Apply filters if needed 4. Click **"Run"** to execute the query 5. View results in the query results panel ## Common tasks ### **Switch between projects** 1. Click the project name in the header 2. Select a project from the dropdown 3. Wait for the project to load and Cube Playground to refresh ### **Sync latest changes** 1. Click **"Sync now"** button in the header 2. Wait for sync to complete 3. Changes from your Git repository will be reflected automatically ### **Add BI integration** 1. Click **"BI Integration"** button 2. Click **"New BI Connection"** 3. Select your BI tool 4. Configure connection settings 5. Sync to activate the connection ## Next steps Now that you have your first project set up: **Manage projects** Learn how to create, edit, and manage multiple Cube projects. **Build cubes** Master the Cube Playground for building and testing cube definitions. **Connect tools** Learn how to use API credentials to connect BI tools and applications. **Sync BI tools** Set up automatic synchronization with your BI tools. ## Tips for success * **Start with Platform Managed** repositories for easier setup * **Use Always On deployment** for production environments * **Protect main branch** in project settings for production work * **Sync regularly** to keep data up-to-date with repository changes * **Document your cubes** with descriptions and examples for team members * **Test queries in Playground** before deploying to production *** **Next: Manage projects** Learn how to create, edit, and configure Cube projects. # Overview Source: https://documentation.5x.co/core-features/metric-store/overview Build and manage metrics data models using Cube.js to define business metrics and enable consistent analytics across your organization **Focus:** Learn how to create and manage metrics data models (Cubes) that provide a unified layer for defining business metrics, enabling consistent analytics across all your tools and applications. The Metrics Store is a comprehensive metrics layer solution built on **Cube.js** technology that enables you to create and manage data models called **Cubes**. These cubes define business metrics and dimensions, providing a standardized way to access your data through SQL and REST APIs, and integrate seamlessly with BI tools. **Powered by Cube.js** The Metrics Store is built on Cube.js, the leading open-source metrics layer platform. We've integrated it deeply with the 5X platform to provide seamless project management, deployment options, and BI tool integration. ## What are cubes? **Cubes** are pre-built data models that contain: * **Measures**: Quantitative metrics like revenue, orders, users, or any calculated business KPIs * **Dimensions**: Attributes for analysis like date, region, product category, or customer segment * **Time dimensions**: Special dimensions for time-based analysis and period-over-period comparisons Cubes provide a standardized way to define business metrics and enable consistent analytics across your organization. Define once, use everywhere. ## Core capabilities ### **Project management** Create and manage multiple Cube projects with Git-based version control: * **Platform managed repositories** - 5X-managed Git repositories for easy setup * **Imported repositories** - Connect your existing GitHub repositories * **Branch management** - Work with multiple branches for different environments or features * **Project settings** - Configure deployment options, protected branches, and environments ### **Cube Playground** Interactive development environment for building and testing cubes: * **Visual query builder** - Drag-and-drop interface for building queries * **SQL editor** - Direct SQL query execution and testing * **Data exploration** - Explore cubes, measures, and dimensions interactively * **Schema management** - Define and edit cube schemas using YAML or JavaScript ### **API access** Query your cubes through multiple interfaces: * **SQL API** - PostgreSQL-compatible connection for BI tools and SQL clients * **REST API** - RESTful API access for application integration * **Real-time calculations** - Compute metrics on-demand with fresh data * **Connection credentials** - Secure, managed credentials for API access ### **BI integration** Automatically sync cube definitions with BI tools: * **Native connectors** - Support for Metabase * **Automatic synchronization** - Keep metric definitions current across tools * **Individual or bulk sync** - Sync specific connections or all at once * **Status tracking** - Monitor sync status and last sync time per connection ### **Deployment options** Configure when your Cube service runs: * **Always on** - Continuous operation for production environments * **On demand** - Start when needed for development or testing * **Scheduled** - Time-based or cron-based scheduling ## Getting started Ready to create your first Cube project? Here's how to get started: ### **Step 1: Navigate to Metrics Store** 1. Go to the **Metrics Store** page (`/metrics-store`) 2. Or click **"Metrics Store"** from the navigation menu ### **Step 2: Create your first project** If no projects exist, you'll see a **"Create project"** button. Otherwise, use the project dropdown and select **"New project"**. **Quick start in 5 minutes** Step-by-step guide to creating your first Cube project and defining your first cubes. ## Core documentation Explore the comprehensive guides to master the Metrics Store: **Create your first project** Quick start guide to creating your first Cube project and accessing the Cube Playground. **Manage projects and branches** Learn how to create, edit, and manage Cube projects with Git-based version control. **Build and test cubes** Master the Cube Playground interface for building queries and defining cube schemas. **Access your data** Learn how to get SQL and REST API credentials for connecting BI tools and applications. **Sync with BI tools** Set up automatic synchronization of cube definitions with your BI tools. **Configure deployment** Understand deployment modes and scheduling options for your Cube service. ## Key benefits ### **Single source of truth** Define business metrics once in Cube schemas and use them consistently across all your analytics tools. Eliminate discrepancies between different reports and dashboards. ### **Flexible development workflow** * **Git-based version control** - Track changes, collaborate, and manage cube definitions * **Branch-based development** - Work on features in separate branches * **Schema as code** - Define cubes using YAML or JavaScript files * **Testing and validation** - Test queries in Cube Playground before deploying ### **Universal access** Make your cubes available through multiple interfaces: * **SQL API** - Connect any SQL-compatible tool or client * **REST API** - Integrate with custom applications * **BI tools** - Native connectors for popular analytics platforms * **Real-time queries** - Compute metrics on-demand with current data ## Integration with core features The Metrics Store enhances all other 5X capabilities: * **Data Warehousing**: Cubes query your warehouse directly for real-time calculations * **Business Intelligence**: BI tools connect to Metrics Store for consistent metrics * **Data & AI Apps**: Applications access standardized metrics through APIs * **Conversational AI**: Natural language queries use Metrics Store definitions * **Orchestration**: Cube projects can be managed through workflow automation ## Common use cases ### **Executive reporting** Ensure leadership sees consistent numbers: * **KPI standardization** across all executive dashboards and reports * **Period-over-period analysis** with consistent time-based calculations * **Goal tracking** comparing actual performance to targets * **Board presentation** materials with verified, consistent metrics ### **Self-service analytics** Empower business users with trusted metrics: * **Metric discovery** helping users find the right data for their questions * **Confidence in results** knowing all metrics follow approved business logic * **Reduced IT burden** fewer requests for custom reports and analysis * **Guided exploration** suggesting relevant metrics and dimensions ### **Cross-functional collaboration** Enable teams to work with shared definitions: * **Sales and marketing alignment** on lead and conversion metrics * **Finance and operations** using the same cost and efficiency calculations * **Product and engineering** sharing user engagement and performance metrics * **Customer success** teams accessing consistent customer health scores *** **Start with quickstart** Get your workspace running and see Metrics Store in action in 15 minutes. # Project management Source: https://documentation.5x.co/core-features/metric-store/project-management Learn how to create, edit, and manage Cube projects with Git-based version control and branch management **Focus:** Master project management in Metrics Store, including creating projects, managing repositories, working with branches, and configuring project settings. Project management in Metrics Store enables you to organize your Cube definitions into separate projects, each with its own Git repository, branches, and deployment configuration. This allows you to manage different environments, features, and teams independently. ## Creating projects ### **Project creation workflow** 1. **Navigate to Metrics Store** * Go to `/metrics-store` page * Or click **"Metrics Store"** from navigation 2. **Access project creation** * If no projects exist: Click **"Create project"** button * If projects exist: Click project name in header → **"New project"** 3. **Fill project details** * Project name (required, unique, max 50 characters) * Repository type (Platform Managed or Imported) * Environment selection * Deployment options * Protected branches (optional) 4. **Add project** * Click **"Add Project"** button * Wait for repository cloning (if imported) * Project will be created and initialized ### **Repository types** Choose **Platform Managed** when you want 5X to handle repository management: **When to use:** * Creating new projects from scratch * Quick setup without Git configuration * Projects that don't require complex Git workflows * Learning and experimentation **What you get:** * 5X-managed Git repository * Automatic repository initialization * Ready-to-use development environment * No GitHub configuration required **Configuration:** * Project name (used to generate repository name) * Environment selection * Deployment options Choose **Imported** when you have an existing GitHub repository: **When to use:** * Importing existing Cube.js projects * Maintaining version control with your team * Complex projects with custom dependencies * Projects shared across multiple environments **Requirements:** * Public or accessible private GitHub repository * Repository must contain valid Cube.js schema files * Deployment key must be added to GitHub repository **Configuration:** * GitHub repository URL * Subdirectory (optional, if Cube files are in a subfolder) * Deployment key (generated and displayed after creation) * Environment selection * Deployment options ### **Project settings** **Project name** * Unique identifier for your project * Maximum 50 characters * Used for display and repository naming * Can be edited after creation **Environment** * Select the environment where your data warehouse is configured * Different environments can have different data sources * Environment selection affects deployment and API access **Deployment options** * **Always on** - Service runs continuously * **On demand** - Service starts when needed * **Scheduled** - Time-based or CRON-based scheduling See [Deployment Options](/core-features/metric-store/deployment-options) for detailed configuration. **Protected branches** * Specify branches that should be protected from direct changes * Protected branches typically require pull requests for changes * Common choices: `main`, `master`, `production` ## Editing projects ### **Edit project details** 1. **Access project settings** * Click **"Settings"** button in header (gear icon) * Or click project name → **"Edit project"** 2. **Modify settings** * Edit project name * Change environment * Update deployment options * Modify protected branches 3. **Save changes** * Click **"Save"** to apply changes * Changes take effect immediately ### **Editable fields** * **Project name** - Update display name * **Environment** - Change deployment environment * **Deployment options** - Modify how service runs * **Protected branches** - Add or remove protected branches ### **Read-only fields** * **GitHub URL** (for imported repos) - Cannot be changed after creation * **Subdirectory** (if set) - Cannot be modified * **Deploy key** (for imported repos) - Displayed for reference ## Deleting projects ### **Delete workflow** 1. **Access project settings** * Go to project settings * Scroll to **"Danger Zone"** section 2. **Delete confirmation** * Click **"Delete Project"** button * Confirm deletion in modal 3. **Post-deletion** * Project and all associated data are permanently removed * Cannot be undone * Repository data is deleted **Permanent deletion** Deleting a project permanently removes: * All cube definitions * All branches and history * All API credentials * All BI connections * All project settings This action cannot be undone. Make sure you have backups of important cube schemas before deleting. ## Branch management ### **Understanding branches** Branches allow you to work on different versions of your cube definitions: * **Main/master branch** - Production-ready cube definitions * **Feature branches** - Development work for new features * **Environment branches** - Different configurations for dev/staging/prod ### **Branch operations** **Switch branches:** 1. Click branch name in header 2. Select branch from dropdown 3. System automatically: * Loads files for selected branch * Updates Cube Playground * Refreshes API credentials **Delete branch:** * Available in project settings * Only non-protected branches can be deleted * Cannot delete currently selected branch ### **Protected branches** **Purpose:** * Prevent accidental changes to critical branches * Enforce code review workflows * Maintain production stability **Configuration:** * Set in project settings * Specify branch names (e.g., `main`, `master`, `production`) * Protected branches cannot be deleted or directly modified ## Repository synchronization ### **Sync now** Manually pull latest changes from Git repository: 1. **Click "Sync now" button** in header 2. **Wait for sync to complete** 3. **Changes reflected automatically:** * Files are refreshed * Cube Playground reloads * API credentials update if needed **When to sync:** * After pushing changes from external Git client * When collaborators make changes * After merging pull requests * To ensure you have latest schema definitions ### **Automatic sync** The system automatically: * Syncs when switching projects * Refreshes when switching branches * Updates when project is first loaded ## Project selection ### **Selecting active project** 1. **Click project name** in header 2. **Select project** from dropdown menu 3. **System automatically:** * Switches to selected project * Starts heartbeat polling * Loads files and branches * Initializes Cube Playground ### **Project list** The dropdown shows: * All available projects * Project names * Current project indicator * **"New project"** option (if you have permissions) ## Best practices ### **Project organization** **Separate environments** Use separate projects for dev, staging, and production to isolate changes and prevent accidents. **Clear project names** Use descriptive names that indicate purpose: "Sales Analytics", "Marketing Metrics", "Finance KPIs". **Protect production** Always protect main/master branches in production projects to prevent accidental changes. **Stay updated** Sync regularly to ensure you're working with the latest cube definitions and changes. ### **Branch management** * **Use descriptive branch names** - `feature/new-metrics`, `fix/revenue-calculation` * **Protect critical branches** - Always protect main/master branches * **Sync before switching** - Ensure you have latest changes before switching branches * **Clean up unused branches** - Delete branches that are no longer needed * **Review before merging** - Use pull requests for protected branches ### **Repository selection** * **Platform Managed** - Best for new projects and quick setup * **Imported** - Best for existing projects and team collaboration * **Consider team workflows** - Choose based on your team's Git practices ## Troubleshooting ### **Common issues** **Possible causes:** * Project name already exists * Environment not selected * Invalid GitHub URL (for imported repos) * Deployment key not added to GitHub **Solutions:** * Use a unique project name * Verify environment is configured * Check GitHub repository accessibility * Ensure deployment key is added to repository settings **Possible causes:** * Branch hasn't been synced from repository * Protected branch restrictions * Repository not fully cloned **Solutions:** * Click "Sync now" to refresh branches * Check project settings for branch restrictions * Wait for repository cloning to complete **Possible causes:** * Git repository access issues * Network connectivity problems * Repository structure issues **Solutions:** * Verify repository access permissions * Check network connection * Review repository structure and files * Try refreshing the page and syncing again *** **Build cubes** Learn how to use Cube Playground to define and test your cubes. **Configure deployment** Understand deployment modes and scheduling options. # Alerting & notifications Source: https://documentation.5x.co/core-features/orchestration/alerting Set up comprehensive alerting for job success, failure, and performance monitoring **Focus:** Learn how to configure intelligent alerting and notifications to stay informed about your job execution status and performance. 5X Jobs provides comprehensive alerting capabilities to keep you informed about job execution status, performance issues, and critical events. Set up notifications via email and Slack to ensure you never miss important job updates. ## Alerting overview ### **Why set up alerts?** **Stay informed:** * Immediate notification of job failures * Success confirmations for critical workflows **Proactive management:** * Address issues before they impact business * Monitor job health and reliability * Track performance trends * Ensure SLA compliance **Team coordination:** * Notify relevant team members * Escalate critical failures * Share success metrics * Coordinate response efforts ## Setting up alerts ### **Creating job alerts** Configure alerts for specific jobs: 1. Navigate to your job in the **Jobs** section 2. Click the **bell icon** or **"no alerts set"** indicator 3. Select **"Set alert"** to configure notifications 4. Choose alert types and notification channels Job Alert Setup Interface ### **Alert configuration options** **Alert types:** * **Job failed** - Notify when job execution fails * **Job success** - Confirm successful job completion **Notification channels:** * **Email** - Direct email notifications * **Slack** - Channel or direct message notifications *** # Datasets & SQL Lab Source: https://documentation.5x.co/core-features/orchestration/execution-monitoring Explore datasets, write custom SQL queries, and manage data sources for powerful analytics **Focus:** Learn how to explore datasets, write custom SQL queries, and leverage SQL Lab for advanced data analysis and visualization. 5X Business Intelligence provides powerful dataset management and SQL Lab capabilities that enable you to explore your data, write custom queries, and create sophisticated analytics. Whether you're a business analyst or data engineer, these tools give you the flexibility to work with data exactly how you need it. ## Dataset management ### **Understanding datasets** Datasets are the foundation of your analytics in 5X Business Intelligence. They represent your data sources and provide the raw material for creating charts, dashboards, and reports. **What datasets provide:** * **Data access** - Direct connection to your warehouse tables and views * **Schema information** - Column types, relationships, and metadata * **Data preview** - Sample data to understand content and structure * **Usage tracking** - Monitor which datasets are used in visualizations * **Security controls** - Row-level security and access permissions ### **Dataset exploration** **Browse available datasets:** Dataset Browser Interface **Key exploration features:** * **Dataset catalog** - Browse all available data sources * **Schema viewer** - Examine table structures and column details * **Data sampling** - Preview actual data to understand content * **Relationship mapping** - See how datasets connect to each other * **Usage analytics** - Track which datasets are most popular ### **Dataset configuration** **Basic dataset settings:** * **Name and description** - Clear identification and business context * **Tags and categories** - Organization for easy discovery * **Ownership** - Assign responsibility for dataset management * **Refresh schedules** - Automatic data updates for real-time insights **Advanced configuration:** * **Custom SQL** - Define calculated fields and transformations * **Caching settings** - Optimize performance with appropriate caching * **Security policies** - Row-level security and access controls * **Data quality rules** - Validation and monitoring settings ### **Dataset importance** **Why datasets matter:** * **Single source of truth** - Centralized access to business data * **Consistency** - Standardized data definitions across the organization * **Efficiency** - Reusable data sources for multiple analyses * **Governance** - Controlled access and usage tracking * **Performance** - Optimized queries and caching strategies ## SQL Lab ### **Introduction to SQL Lab** SQL Lab is 5X Business Intelligence's powerful query interface that allows you to write and execute custom SQL queries directly against your data warehouse. **Key capabilities:** * **Full SQL support** - Complex queries with joins, subqueries, and advanced functions * **Query history** - Track and reuse previous queries * **Query sharing** - Collaborate with team members * **Result export** - Export query results in various formats * **Query optimization** - Built-in performance suggestions SQL Lab Query Interface ### **Writing effective SQL queries** **Query structure best practices:** ```sql theme={null} -- Example: Well-structured query with clear formatting SELECT DATE_TRUNC('month', order_date) AS month, region, COUNT(*) AS order_count, SUM(order_amount) AS total_revenue, AVG(order_amount) AS avg_order_value FROM orders WHERE order_date >= '2024-01-01' AND order_status = 'completed' GROUP BY 1, 2 ORDER BY month DESC, total_revenue DESC LIMIT 100; ``` **Query optimization techniques:** * **Use appropriate filters** - Limit data with WHERE clauses * **Select only needed columns** - Avoid SELECT \* for better performance * **Leverage indexes** - Structure queries to use database indexes * **Limit result sets** - Use LIMIT to control output size * **Use aggregations** - Pre-aggregate data when possible ### **Advanced SQL features** **Window functions for analytics:** ```sql theme={null} -- Example: Using window functions for advanced analytics SELECT customer_id, order_date, order_amount, SUM(order_amount) OVER ( PARTITION BY customer_id ORDER BY order_date ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW ) AS running_total, ROW_NUMBER() OVER ( PARTITION BY customer_id ORDER BY order_date DESC ) AS order_rank FROM orders; ``` **Common table expressions (CTEs):** ```sql theme={null} -- Example: Using CTEs for complex analysis WITH monthly_sales AS ( SELECT DATE_TRUNC('month', order_date) AS month, SUM(order_amount) AS total_sales FROM orders GROUP BY 1 ), sales_growth AS ( SELECT month, total_sales, LAG(total_sales) OVER (ORDER BY month) AS prev_month_sales, (total_sales - LAG(total_sales) OVER (ORDER BY month)) / LAG(total_sales) OVER (ORDER BY month) * 100 AS growth_rate FROM monthly_sales ) SELECT * FROM sales_growth; ``` ### **SQL Lab features** **Query management:** * **Save queries** - Store frequently used queries for reuse * **Query templates** - Create reusable query patterns * **Version control** - Track changes to saved queries * **Collaboration** - Share queries with team members **Result handling:** * **Export options** - CSV, Excel, JSON formats * **Result caching** - Store query results for faster access * **Large result sets** - Handle big data efficiently * **Visualization** - Quick chart creation from query results ## Creating charts from SQL ### **SQL as data source** You can use custom SQL queries as the data source for charts, providing unlimited flexibility: **When to use SQL for charts:** * **Complex calculations** - Multi-step data transformations * **Advanced filtering** - Sophisticated WHERE clauses * **Data joins** - Combining multiple data sources * **Performance optimization** - Optimized queries for specific use cases **SQL chart workflow:** 1. **Write your query** - Use SQL Lab to create the query 2. **Test and validate** - Execute and verify results 3. **Create chart** - Use query results as data source 4. **Configure visualization** - Choose appropriate chart type 5. **Add to dashboard** - Include in your dashboard layout ### **Parameterized queries** Use dashboard filters to make your SQL queries dynamic: ```sql theme={null} -- Example: Parameterized query using dashboard filters SELECT DATE_TRUNC('{{ granularity }}', order_date) AS time_period, region, SUM(order_amount) AS total_revenue FROM orders WHERE order_date >= '{{ start_date }}' AND order_date <= '{{ end_date }}' {% if region %} AND region = '{{ region }}' {% endif %} GROUP BY 1, 2 ORDER BY time_period DESC; ``` **Parameter types:** * **Date parameters** - Dynamic date ranges * **Text parameters** - Categorical filters * **Numeric parameters** - Value range filters * **Boolean parameters** - True/false conditions ## Data quality and validation ### **Data profiling** **Understanding your data:** * **Completeness** - Percentage of non-null values * **Uniqueness** - Duplicate record identification * **Consistency** - Data format and value consistency * **Accuracy** - Data correctness and validity * **Timeliness** - Data freshness and update frequency **Profiling techniques:** * **Column analysis** - Data types, ranges, and distributions * **Pattern recognition** - Identify common data patterns * **Outlier detection** - Find unusual or suspicious values * **Relationship analysis** - Understand data dependencies ### **Data validation queries** ```sql theme={null} -- Example: Data quality validation queries -- Check for missing values SELECT COUNT(*) AS total_records, COUNT(customer_id) AS non_null_customer_ids, COUNT(*) - COUNT(customer_id) AS missing_customer_ids FROM orders; -- Check for duplicate records SELECT customer_id, order_date, COUNT(*) AS duplicate_count FROM orders GROUP BY 1, 2 HAVING COUNT(*) > 1; -- Check for outliers SELECT customer_id, order_amount FROM orders WHERE order_amount > ( SELECT AVG(order_amount) + 3 * STDDEV(order_amount) FROM orders ); ``` ## Best practices ### **Dataset management** **Clear structure** Use consistent naming conventions, tags, and categories to organize your datasets effectively. **Add context** Include descriptions, business context, and usage notes for each dataset. **Track performance** Monitor dataset usage patterns and optimize frequently accessed datasets. **Control access** Implement appropriate security policies and row-level security where needed. ### **SQL development** **Build incrementally** Start with simple queries and gradually add complexity as you understand the data better. **Add comments** Include comments in your SQL queries to explain complex logic and business rules. **Validate results** Always test queries with small datasets before running on large data volumes. **Efficient queries** Write efficient queries that perform well and don't impact system performance. ## Troubleshooting ### **Common issues** **Possible causes:** * Insufficient permissions for dataset access * App connection configuration problems * Data source connectivity issues * Row-level security restrictions **Solutions:** * Verify user permissions and roles * Check App Connection settings * Test data source connectivity * Review security policies **Possible causes:** * Inefficient query structure * Large dataset volumes * Missing database indexes * Complex calculations or joins **Solutions:** * Optimize query structure and logic * Use appropriate filters and limits * Check database indexing strategy * Consider data aggregation approaches **Possible causes:** * Missing or null values * Data format inconsistencies * Outdated or stale data * Data integration problems **Solutions:** * Implement data validation checks * Standardize data formats * Update data refresh schedules * Review data integration processes *** **Create visualizations** Use your datasets and SQL queries to build compelling charts and dashboards. **Build dashboards** Combine your datasets and SQL insights into interactive dashboards. **Connect data sources** Learn how to set up App Connections to access your data warehouse. **Explore your data** Dive deeper into data exploration techniques and advanced analytics. *** # Overview Source: https://documentation.5x.co/core-features/orchestration/overview Visual workflow automation with dependency tracking and monitoring - orchestrate your data pipelines with ease **Focus:** Understand how 5X Jobs automates your data workflows with visual pipeline building, flexible scheduling, and comprehensive monitoring. 5X Jobs is a visual workflow automation platform that automates your data workflows through drag-and-drop pipeline building. Create complex data pipelines with triggers, tasks, and dependencies - all managed through an intuitive no-code interface that handles scheduling, monitoring, and execution automatically. 5X Jobs Dashboard - Empty State ## Key capabilities ### **Visual workflow builder** Create complex data pipelines using an intuitive drag-and-drop interface: * **Trigger nodes** - Start workflows with manual, scheduled, webhook, or pull request triggers * **Task nodes** - Execute data ingestion, dbt modeling, custom Python code, or parallel jobs * **Dependency management** - Define execution order and parallel processing with visual connections * **Real-time preview** - See your pipeline structure and flow before execution ### **Flexible scheduling options** Choose how your jobs are triggered based on your workflow needs: **On-demand execution** Run jobs instantly with a single click for immediate processing. **Use cases:** * Testing and debugging workflows * Ad-hoc data processing * Emergency data updates * Development and experimentation **Time-based execution** Automate jobs with precise scheduling intervals and timezone support. **Features:** * Intervals as frequent as 1 minute * Daily, weekly, monthly schedules * Custom CRON expressions * Automatic timezone handling **Event-driven execution** Trigger jobs from external systems with secure authentication. **Features:** * Real-time pipeline activation * Configurable access tokens * Secure authentication * Integration with external APIs **GitHub CI/CD integration** Automatically execute jobs when pull requests are created or updated. **Features:** * GitHub integration * Automatic job execution on PR creation * Draft pull request support * CI/CD workflow automation ### **Comprehensive task automation** Execute different types of work within your pipelines: **Data ingestion:** * Use the same 600+ connectors available in 5X Ingestion * Seamlessly integrate data collection into your workflows * Automatic dependency management with other pipeline steps **dbt modeling:** * Execute dbt runs with configurable commands and parameters * Support for multiple dbt versions (1.6.18, 1.7.19, 1.8.9, 1.9.10) * Support for multiple repositories and branches * Integration with dbt documentation generation * Deferral support for comparing changes across environments **Custom Python code:** * Run Python scripts from GitHub repositories * Support for Python versions 3.8.19 through 3.12.4 * Dependency management with requirements files * Environment variable and secret injection **Parallel job execution:** * Execute other jobs as sub-workflows * Enable complex pipeline orchestration and reuse * Hierarchical job management across your workspace ### **Real-time monitoring and management** Track job execution with comprehensive visibility: * **Job run history** - View all past executions with status, duration, and resource usage * **Real-time status** - Monitor active jobs with live updates * **Detailed logs** - Access execution logs for debugging and troubleshooting * **Performance metrics** - Track compute usage, execution time, and resource consumption * **Draft management** - Develop and test changes before publishing to production 5X Jobs Monitoring Dashboard ### **Intelligent alerting system** Stay informed about your job execution with comprehensive notifications: * **Success/failure alerts** - Get notified when jobs complete or fail * **Multiple channels** - Email and Slack integration * **Granular control** - Set alerts at job level or workspace level * **Real-time delivery** - Immediate notifications for critical events ### **Environment and resource management** Run jobs across different environments with proper isolation: * **Environment configuration** - Define separate environments for development, staging, and production * **App connection integration** - Link environments to specific warehouse connections * **Environment variables** - Manage secrets and configuration across environments * **Compute profiles** - Configure resource allocation per job execution * **Concurrent execution control** - Limit parallel job runs to manage resource usage ## Getting started Ready to automate your first data workflow? Here's how to get started: **Set up app connections** Configure your data warehouse connection and environment before creating jobs. **Build a workflow** Learn how to create triggers, add tasks, and build your first automated pipeline. *** # Workflow builder Source: https://documentation.5x.co/core-features/orchestration/workflow-builder Learn how to create visual data pipelines with triggers, tasks, and dependencies **Focus:** Master the visual workflow builder to create automated data pipelines with triggers, tasks, and intelligent dependency management. The 5X Jobs workflow builder provides an intuitive visual interface for creating complex data pipelines. Build workflows by connecting triggers and tasks, define execution dependencies, and configure each component to meet your specific requirements. ## Creating a new job ### **Step 1: Name your job** Start by creating a new job and giving it a descriptive name: 1. Navigate to the **Jobs** section in your workspace 2. Click **"+ New Job"** to create a new workflow 3. Enter a descriptive name for your job in the modal dialog 4. Click **"Create job"** to proceed to the workflow builder ### **Step 2: Configure job settings** Before building your workflow, configure the basic job settings: * **Job name** - Update the job name if needed * **Compute profile** - Select the appropriate compute resources (Small, Medium, Large) * **Environment** - Choose the environment where this job will execute * **Timeout settings** - Set execution time limits if needed * **Concurrent runs** - Configure how many instances can run simultaneously Job Settings Configuration ## Building workflows with triggers Every job workflow starts with a trigger that determines when and how the pipeline executes. Perfect for testing, debugging, or ad-hoc execution: 1. Click **"Start the flow"** on the canvas 2. Select **"Manual"** from the trigger options 3. Configure the trigger name 4. Click to add the trigger to your workflow Manual Trigger Configuration **Use cases:** * Testing new workflows * Ad-hoc data processing * Manual data refresh * Debugging pipeline issues Automate execution based on time intervals: 1. Select **"Scheduled"** from the trigger options 2. Configure the trigger name 3. Choose scheduling method: * **Custom Time** - Set intervals (every minute, hour, day, week, month) * **CRON Expression** - Use advanced scheduling syntax 4. Set the execution time and timezone Scheduled Trigger Configuration **Scheduling options:** * **Frequency** - Every minute, hour, day, week, or month * **Time** - Specific time of day for execution * **Timezone** - Automatic UTC handling with local time display * **CRON** - Advanced scheduling for complex patterns Enable event-driven execution from external systems: 1. Select **"Webhook"** from the trigger options 2. Configure the trigger name 3. Enable authorization (recommended for security) 4. Save to generate webhook URL and access token Webhook Trigger Configuration **Features:** * **Secure authentication** - Configurable access tokens * **Real-time execution** - Immediate pipeline activation * **External integration** - Connect with any system that can send HTTP requests Integrate with GitHub for CI/CD workflows: 1. Select **"Pull Request"** from the trigger options 2. Configure the trigger name 3. Select the GitHub repository 4. Optionally enable **"Run on draft pull request"** 5. Note the access token requirements Pull Request Trigger Configuration **Requirements:** * GitHub repository access * Contact 5X support for access token * Configure `DBT_TOKEN` in GitHub Secrets ## Adding tasks to your workflow Once you have a trigger, add tasks to define what work your pipeline performs. Use existing connectors to bring data into your warehouse: 1. Click the **"+"** button after your trigger 2. Select **"Ingestion"** from the task options 3. Configure the task: * **Task name** - Descriptive name for this ingestion step * **Connector** - Select from 600+ available connectors * **Depends on** - Choose trigger or previous tasks Ingestion Task Configuration **Integration benefits:** * Same connectors as 5X Ingestion * Automatic dependency management * Seamless data flow between pipeline steps Execute dbt transformations and models: 1. Select **"Modeling"** from the task options 2. Configure the modeling task: * **Task name** - Name for this modeling step * **Repository** - GitHub repository containing dbt project * **Branch** - Specific branch to use * **Depends on** - Choose trigger or previous tasks * **dbt commands** - Commands to execute (default: `dbt run`) * **Target name** - dbt target environment * **Number of threads** - Parallel execution threads * **Deferral** - Compare against another environment * **Generate dbt docs** - Automatically create documentation dbt Modeling Task Configuration **dbt integration features:** * Multiple dbt versions supported (1.6.18, 1.7.19, 1.8.9, 1.9.10) * Custom command execution * Environment-specific targeting * Documentation generation * Deferral for change comparison Execute Python scripts and applications: 1. Select **"Custom Code"** from the task options 2. Configure the Python task: * **Task name** - Name for this code execution * **Repository** - GitHub repository with Python code * **Branch** - Specific branch to use * **Python file** - Main script to execute (by name or path) * **Dependent libraries** - Requirements file (by name or path) * **Python version** - Select from supported versions (3.8.19 - 3.12.4) * **Depends on** - Previous tasks or triggers Custom Python Code Task Configuration **Python execution features:** * Multiple Python versions (3.8.19, 3.9.19, 3.10.14, 3.11.9, 3.12.4) * Dependency management with requirements files * Environment variable injection * Secret management for sensitive data Execute other jobs as sub-workflows: 1. Select **"Job"** from the task options 2. Configure the job task: * **Task name** - Name for this sub-job execution * **Job** - Select another job from your workspace * **Depends on** - Previous tasks or triggers Parallel Job Task Configuration **Parallel execution benefits:** * Reuse existing job definitions * Create hierarchical workflows * Enable complex orchestration patterns * Modular pipeline design ## Managing dependencies and flow ### **Visual dependency management** The workflow builder automatically handles task dependencies: * **Sequential execution** - Tasks run in order based on connections * **Parallel execution** - Tasks without dependencies run simultaneously * **Conditional flow** - Complex branching based on task outcomes * **Error handling** - Failed tasks stop dependent execution Complete Workflow with Tasks and Dependencies ### **Best practices for workflow design** **Start simple:** * Begin with basic triggers and single tasks * Test each component before adding complexity * Use manual triggers for initial development **Design for reliability:** * Add appropriate timeout settings * Configure concurrent run limits * Plan for error scenarios and retries **Optimize for performance:** * Use parallel execution where possible * Select appropriate compute profiles * Minimize unnecessary dependencies **Maintain clarity:** * Use descriptive names for triggers and tasks * Document complex workflows * Keep workflows focused on specific business processes ## Publishing and testing workflows ### **Draft vs Live versions** 5X Jobs supports draft development with live execution: * **Draft changes** - Develop and modify workflows without affecting production. Test changes thoroughly before publishing! * **Live execution** - Only published workflows execute on schedule ### **Publishing workflows** 1. Complete your workflow configuration 2. Review all settings and dependencies 3. Click **"Publish"** to make changes live 4. Monitor the first execution to ensure proper operation *** # Welcome to 5X Source: https://documentation.5x.co/index The end-to-end data and AI platform that eliminates data fragmentation and accelerates intelligent applications Hero Light Hero Dark ## Introduction 5X is the end-to-end data and AI platform that helps organizations eliminate data fragmentation, streamline analytics, and rapidly build intelligent applications. Our platform integrates your disconnected data sources and transforms them into actionable insights, through dashboards, self-serve analytics, and AI-powered apps. > 5X accelerates your data foundation so you can unlock real-world AI use cases, fast and reliably. ## Why 5X? Businesses today face critical challenges: * **Siloed, messy data** across departments * **Low-quality, unreliable pipelines** causing mistrust * **Slow delivery cycles** for business teams and data teams * **A desire to adopt AI**, but without the right foundations 5X was built to solve these exact problems. We unify modern data infrastructure in a single platform so you don't have to stitch it together yourself. ## Get Started Ready to transform your data journey? Start with these essential guides to deploy and maximize 5X. Deploy 5X in your environment and connect your first data sources ## Built For Every Team Whether you're a business user seeking insights or a data engineer building pipelines, 5X delivers tailored solutions for your needs. Access consistent metrics, self-service analytics, and custom AI applications 600+ connectors, automated orchestration, and governance out-of-the-box ## Platform Capabilities Explore the comprehensive features that make 5X the complete data and AI solution. Ingestion, Reverse ETL, and data transport capabilities Dashboards, BI tools, GenAI & predictive analytics, custom data apps 600+ out-of-the-box connectors to integrate all your data sources Automated orchestration, metrics layers, and data modeling Built-in lineage, access controls, and data governance features Flexible deployment: 5X SaaS cloud or private cloud (AWS, Azure, etc.) # Overview Source: https://documentation.5x.co/quickstart Get your workspace running and querying real data with your team in under 15 minutes **Focus:** Get up and running quickly with immediate value, not deep education. This streamlined guide gets you from zero to querying real data with your team in about 15 minutes. You'll walk away with a working environment and clear signal that the 5X platform delivers value. **Estimated Total Time:** 15 minutes This activation-focused quickstart covers only the essentials. Advanced features like modeling, dashboards, and orchestration are available when you're ready to go deeper. ## Essential quickstart steps Follow these steps to get your workspace operational: Complete your initial setup from account creation through workspace provisioning. **Time:** \~5 minutes **Time:** \~2 minutes Set up secure connections to your data sources. **Time:** \~3 minutes Connect and sync data from your sources into your warehouse. **Time:** \~3 minutes **Critical milestone:** Query and explore your real ingested data using SQL. **Time:** \~2 minutes Add teammates and set up basic permissions for collaboration. **Optional:** Explore advanced capabilities when ready Learn about modeling, dashboards, orchestration, and more. ## What this quickstart accomplishes After completing steps 1-7, you'll have: ✅ **Working data workspace** with your warehouse connected\ ✅ **Real data flowing** from your sources\ ✅ **SQL environment ready** for immediate data exploration\ ✅ **Team collaboration set up** with appropriate permissions\ ✅ **Foundation in place** for all advanced features ## What this quickstart intentionally skips These powerful features require context and a learning mindset - they're covered in Step 8 and dedicated guides when you're ready: ❌ **dbt Modeling** - Requires understanding of models, repos, dependencies\ ❌ **Orchestration (Jobs)** - Introduces complex scheduling and dependency logic\ ❌ **BI Dashboards** - Needs metrics, modeling, and stakeholder clarity\ ❌ **Metrics Store** - Requires data modeling and metric definitions\ ❌ **Data & AI Apps** - Needs use case clarity and data preparation\ ❌ **Advanced Git/CI-CD** - Better suited for mature implementations ## Ready to start? **Start here:** Account setup and workspace provisioning Get your 5X workspace up and running in minutes. *** ## Alternative paths Already have a workspace? Jump to advanced capabilities. # Steps 1-3: Account Setup and Workspace Provisioning Source: https://documentation.5x.co/quickstart/step-1-create-account Complete your initial setup from account creation to workspace provisioning **Estimated Time:** 10-12 minutes Complete your initial setup from account creation through workspace provisioning in one comprehensive workflow. ## Setup overview This combined setup process will take you through three essential steps: 1. **Create your 5X account** (\~2 minutes) 2. **Personal and workspace setup** (\~2 minutes) 3. **Configure data warehouse and workspace provisioning** (\~5-7 minutes) *** ## Step 1: Create your 5X account **Account access:** 5X provides private signup links to new customers. You'll receive a personalized signup link from the 5X team to create your account. **Need a signup link?** Contact your 5X sales representative or reach out to [support@5x.co](mailto:support@5x.co) to get started with your new account and workspace setup. ### Account creation options The 5X Platform offers two convenient ways to create your account using your business email through the private signup link: **Best for:** Direct registration with corporate credentials * Use your business email address * Create a secure password * Email verification required **Best for:** Quick setup with existing Google Workspace accounts * One-click authentication with Google * Use your business Google account * Automatic email verification * Secure OAuth integration ### Account creation process 1. **Access the private signup link** * Use the private signup link provided by the 5X team * You'll see the "Sign up on 5X" page with registration options 5X Platform signup page showing account creation options with Google authentication and email signup 2. **Choose your registration method** **For Google authentication:** * Click **"Sign up with Google"** button * Select your business Google account * Grant necessary permissions * Email automatically verified **For corporate email:** * Click **"Or sign up with email"** * Enter your work email address in the "Work Email" field * Click **"Sign up with email"** button 3. **Email verification** (corporate email only) * Check your inbox for verification email * Click the verification link. You'll be redirected back to the platform. Complete the signup process. *** ## Step 2: Personal and workspace setup Once your account is verified, you'll complete two quick setup forms. ### Personal information Complete your profile with the "Let's get to know you" form: 1. **Enter your name** in the "What should we call you?" field 2. **Set a secure password** that meets the displayed requirements 3. Click **"Continue"** when all checkmarks are green ### Workspace setup Configure your workspace with the "Setup your first workspace" form: Workspace setup form showing workspace name field and warehouse selection options 1. **Enter a workspace name** (e.g., "Acme Corp Analytics") 2. **Choose your warehouse option:** * **"No, I don't have a warehouse yet"** - 5X will create one for you (recommended) * **"Yes, I already have an existing warehouse"** - Connect your existing warehouse 3. **Add referral code** (optional) 4. Click **"Continue with 5X provisioned warehouse"** *** ## Step 3: Configure data warehouse and workspace provisioning Choose your data storage solution based on your organization's needs. After configuration, your workspace will be automatically provisioned. ### Warehouse options **Recommended for quick evaluations and for users who don't yet have a warehouse set up** * Fully managed and optimized * No setup complexity * Automatic scaling * Built-in security **For organizations with existing infrastructure** * Connect Snowflake, Google BigQuery, AWS Redshift, or PostgreSQL * Maintain existing investments * Custom configurations ### Option A: 5X managed warehouse setup 1. **Configure your warehouse settings** * You'll see the "Setup your warehouse" screen 5X managed warehouse setup screen showing plan, cloud, and region selection with free trial offer 2. **Select your configuration** * **Plan:** Choose from available plans (Standard is pre-selected) * **Cloud:** Select your preferred cloud provider (AWS, GCP, or Azure) * **Region:** Choose the region closest to your users for optimal performance 3. **Complete setup and begin provisioning** * Review your selections * Click **"Continue with free trial"** to proceed * Your workspace with a 5X managed snowflake warehouse will be provisioned. The provisioning process will take around 5 minutes. ### Option B: Connect existing warehouse 1. **Select your warehouse type** Enterprise cloud data warehouse Google Cloud serverless warehouse Amazon cloud data warehouse Open source relational database 2. **Enter connection details** **Step 1: Prerequisites** Before proceeding, ensure you have: 1. **Required Snowflake roles:** ACCOUNTADMIN, SECURITYADMIN, and SYSADMIN 2. **Network access:** If you have network policies on your Snowflake account, allowlist these static IPs: ``` 35.236.200.239/32 35.230.162.249/32 ``` **Step 2: Create Snowflake user & role for 5X** Create a SQL file and run the following script in your Snowflake console: ```sql theme={null} begin; -- create variables for user / password / role / warehouse / database (needs to be uppercase for objects) set role_name = 'FIVEX_ROLE'; set user_name = 'FIVEX_USER'; set user_password = 'yourpassword'; set warehouse_name = 'FIVEX_WAREHOUSE'; -- change role to securityadmin for user / role steps use role securityadmin; -- create role for fivex create role if not exists identifier($role_name); grant role identifier($role_name) to role SYSADMIN; -- create a user for fivex create user if not exists identifier($user_name) password = $user_password default_role = $role_name default_warehouse = $warehouse_name; grant role identifier($role_name) to user identifier($user_name); -- set binary_input_format to BASE64 ALTER USER identifier($user_name) SET BINARY_INPUT_FORMAT = 'BASE64'; -- change role to sysadmin for warehouse / database steps use role sysadmin; -- create a warehouse for fivex (optional) create warehouse if not exists identifier($warehouse_name) warehouse_size = xsmall warehouse_type = standard auto_suspend = 60 auto_resume = true initially_suspended = true; -- grant fivex role access to warehouse grant USAGE on warehouse identifier($warehouse_name) to role identifier($role_name); --grant fivex access to users and roles for team management use role ACCOUNTADMIN; grant create user on account to role identifier($role_name); grant create role on account to role identifier($role_name); grant manage grants on account to role identifier($role_name); -- change role to ACCOUNTADMIN for utilization & audit logs related steps grant IMPORTED PRIVILEGES on database SNOWFLAKE to role identifier($role_name); use role sysadmin; commit; ``` **Important:** Replace 'yourpassword' in the script above with a strong, secure password. You can customize names, but we recommend keeping FIVEX\_USER, FIVEX\_ROLE and FIVEX\_WAREHOUSE. **Step 3: Upload authentication key pair** For secure key-pair authentication: 1. Generate a private/public key pair (if not already done) 2. In the 5X platform, enter: * **Account URL:** `https://account_name.region.snowflakecomputing.com` * **Username:** `FIVEX_USER` (or your custom username) * **Private Key:** Paste your private key content * **Passphrase:** Only required if your private key is encrypted 3. Run this script to update the RSA public key in Snowflake: ```sql theme={null} ALTER USER SET RSA_PUBLIC_KEY = ''; ``` **Step 4: Enter connection details** Complete the connection form with your existing Snowflake configuration: * **Default Warehouse:** `FIVEX_WAREHOUSE` (or your custom warehouse name) * **Default Role:** `FIVEX_ROLE` (or your custom role name) * **Plan:** Enter your current Snowflake plan/edition * **Cloud:** Specify your existing cloud provider (AWS, GCP, or Azure) * **Region:** Enter the region where your Snowflake account is hosted **Step 1: Pre-requisites** 1. Log in to AWS Console 2. Navigate to Amazon Redshift 3. Select your cluster and copy the Endpoint URL (without jdbc:redshift:// prefix) ``` http://cluster-name.account-id.region.redshift.amazonaws.com ``` 4. Database name is the initial database created with your cluster. You can find this information in your cluster configuration. 5. Use username and password authentication for database access. Ensure the user has appropriate permissions for your use case. 6. (Optional) Run on Amazon Redshift SQL editor for connection details. ```sql theme={null} begin; -- Create variables (for reference; Redshift doesn’t support SQL variables directly) -- user_name = 'FIVEX_USER'; -- user_password = 'yourpassword'; -- database_name = 'FIVEX_DATABASE'; -- schema_name = 'FIVEX_SCHEMA'; -- Create a user CREATE USER FIVEX_USER PASSWORD 'yourpassword'; -- Create a database (optional) CREATE DATABASE FIVEX_DATABASE; -- Grant DB-level permissions GRANT CREATE ON DATABASE FIVEX_DATABASE TO FIVEX_USER; GRANT TEMPORARY ON DATABASE FIVEX_DATABASE TO FIVEX_USER; -- Create schema CREATE SCHEMA IF NOT EXISTS FIVEX_SCHEMA; -- Grant schema-level permissions GRANT ALL ON SCHEMA FIVEX_SCHEMA TO FIVEX_USER; GRANT ALL PRIVILEGES ON ALL TABLES IN SCHEMA FIVEX_SCHEMA TO FIVEX_USER; -- Grant access to system catalogs GRANT USAGE ON SCHEMA information_schema TO FIVEX_USER; GRANT SELECT ON ALL TABLES IN SCHEMA information_schema TO FIVEX_USER; GRANT USAGE ON SCHEMA pg_catalog TO FIVEX_USER; GRANT SELECT ON ALL TABLES IN SCHEMA pg_catalog TO FIVEX_USER; -- Grant access to all schemas GRANT USAGE ON ALL SCHEMAS IN DATABASE FIVEX_DATABASE TO FIVEX_USER; GRANT SELECT ON ALL TABLES IN DATABASE FIVEX_DATABASE TO FIVEX_USER; -- Set default privileges in FIVEX schema ALTER DEFAULT PRIVILEGES IN SCHEMA FIVEX_SCHEMA GRANT ALL ON TABLES TO FIVEX_USER; -- Set default privileges in public schema ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT ON TABLES TO FIVEX_USER; -- Grant execute on functions GRANT EXECUTE ON ALL FUNCTIONS IN SCHEMA public TO FIVEX_USER; GRANT EXECUTE ON ALL FUNCTIONS IN SCHEMA FIVEX_SCHEMA TO FIVEX_USER; commit; ``` **Step 2: Basic connection information** 1. Enter 'Host' & 'Port' details 2. Enter 'Database' name 3. Enter 'Username' & 'Password' 4. Select 'Cluster Region' 5. Click on 'Next step' to proceed **Step 3: Security and Network configuration** 1. Go to AWS Console > VPC > Security Groups 2. Find your Redshift cluster's security group and Edit inbound rules 3. Add a new rule for each IP with Port 5439 ``` 35.236.200.239/32 35.230.162.249/32 ``` 4. Database name is the initial database created with your cluster. You can find this information in your cluster configuration. 5. If your cluster is in a VPC, also ensure: * Subnet is properly configured * Route tables allow internet access * DNS resolution is enabled **Step 1: Pre-requisites** 1. Open port 5432 (or your custom port) in your firewall to allow inbound traffic from 5X IP addresses 2. Allow 5X's static IP addresses through your firewall or security group ``` 35.236.200.239/32 35.230.162.249/32 ``` 3. The database must be accessible over the internet (or VPN/SSH tunnel) 4. You must have SUPERUSER or database owner privileges. 5. (Optional) Create a 5X user with SUPERUSER permission ```sql theme={null} -- PostgreSQL Setup for 5X Platform BEGIN; -- Create user for 5X platform CREATE USER FIVEX_USER WITH PASSWORD 'your_secure_password'; -- Grant database connection privileges GRANT CONNECT ON DATABASE your_database TO FIVEX_USER; -- Grant schema usage and creation privileges GRANT USAGE, CREATE ON SCHEMA public TO FIVEX_USER; -- Grant table privileges GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA public TO FIVEX_USER; ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO FIVEX_USER; -- Grant sequence privileges GRANT USAGE, SELECT ON ALL SEQUENCES IN SCHEMA public TO FIVEX_USER; ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO FIVEX_USER; -- Grant function privileges GRANT EXECUTE ON ALL FUNCTIONS IN SCHEMA public TO FIVEX_USER; ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT EXECUTE ON FUNCTIONS TO FIVEX_USER; -- Additional privileges for metadata access GRANT SELECT ON pg_catalog.pg_tables TO FIVEX_USER; GRANT SELECT ON information_schema.tables TO FIVEX_USER; GRANT SELECT ON information_schema.columns TO FIVEX_USER; COMMIT; ``` **Step 2: Database connection details** Enter your Postgres connection information 1. Enter 'Host' & 'Port' details 2. Enter 'Database' name 3. Enter 'Username' & 'Password' 4. Click on 'Connect' to finish the process **Step 1 of 4: Enable required APIs in Google Cloud** 1. Go to the [Google Cloud Console](https://console.cloud.google.com) 2. Search for 'APIs & Services' and open the APIs & Services page. Click on 'Enabled APIs & services' button 3. Search and enable: * BigQuery API * Cloud Resource Manager API Ensure both APIs show as "Enabled" in your dashboard. **Important:** Use an account with Project Admin permissions. API changes might take a few minutes to take effect. **Step 2 of 4: Create a Service Account** 1. In the GCP Console, search for 'Service Accounts' and click on Service Accounts 2. Click '+ Create Service Account' 3. Fill in the Service account name and add a description if you'd like (For example: 5X\_Service\_Account) 4. Click 'Create and continue' 5. In the second step, 'Grant this service account access to project', give the following permissions: * BigQuery User * Project IAM Admin (Optional) 'Project IAM Admin' permission enables us to give you an ability to manage Google BigQuery users on the 5X teams page. Click 'Continue' and then click 'Done'. **Step 3 of 4: Generate JSON Key** 1. Go to IAM & Admin > Service Accounts, and click on the service account you just created 2. Go to the 'Keys' tab 3. Click "Add Key" > "Create new key" 4. Choose JSON and click 'Create' 5. Save the downloaded .json file **Important:** Keep the downloaded JSON file secure—it contains sensitive credentials. **Step 4 of 4: Final step** Complete the connection in the 5X platform: 1. Upload the .json file you just downloaded 2. Select all the project ID(s) you want to add to 5X 3. Complete the connection process 3. **Test and validate connection** * Click **"Test Connection"** * Resolve any connection issues * Confirm successful connection * **Workspace provisioning begins automatically** after successful connection test ### Alternative: Book a free setup call If you need assistance with configuring your existing warehouse connection, our team is here to help: **Get personalized assistance with:** * Warehouse connection configuration * Permission setup and troubleshooting * Best practices for your specific environment * Security and performance optimization Our technical team will guide you through the entire setup process to ensure a smooth integration with your existing data infrastructure. **To schedule your call:** 1. Click **"Book a free setup call"** button on the warehouse selection screen 2. Choose a convenient time slot from our calendar 3. Provide brief details about your current warehouse setup 4. Receive calendar invitation with meeting details ## Troubleshooting common issues **Email verification problems:** * Check spam/junk folder * Wait 5-10 minutes for delivery * Request new verification email **Password requirements:** * Minimum 8 characters * Include uppercase, lowercase, number, and special character **Corporate email issues:** * Ensure business domain (not personal email) * Contact IT if domain is blocked * Try Google authentication alternative **Network problems:** * Check firewall settings and IP whitelisting * Verify VPN connectivity * Ensure correct ports are open **Authentication failures:** * Verify credentials in source system * Check service account permissions * Ensure credentials haven't expired **Permission errors:** * Review required permissions checklist * Contact database administrator * Verify role assignments **Stuck provisioning:** * Wait additional 5 minutes * Refresh page for updated status * Check detailed logs for errors **Failed provisioning:** * Verify warehouse connection stability * Check cloud provider service status * Contact [support@5x.co](mailto:support@5x.co) ## What's next? Congratulations! You've successfully completed the initial setup of your 5X Platform workspace. Your environment is now ready for data integration and analytics. **Next:** Configure app connections and credentials Set up secure connections to your data sources and applications. *** Return to quickstart guide overview Contact our support team # Step 4: Configure app connections and credentials Source: https://documentation.5x.co/quickstart/step-4-app-connections-credentials Learn how to configure app connections and credentials to enable seamless integration between 5X capabilities and your data warehouse **Estimated Time:** 5-10 minutes Learn how to configure app connections and credentials to enable seamless integration between 5X capabilities and your data warehouse. **Note for 5X managed warehouse users:** If you provisioned your workspace with a 5X managed warehouse, default app connections and credentials are automatically configured for all capabilities. You can skip this step and return later if you need to modify, delete, or create additional app connections. ## Understanding app connections and credentials App connections and credentials are fundamental concepts that enable the 5X Platform's capabilities to securely connect and interact with your data warehouse. ## What are app connections? App connections define how each 5X capability (ingestion, modeling, orchestration, BI, and metrics layer) connects to your data warehouse. Think of them as capability-specific connection profiles that contain all the technical details needed for that particular workload to access your warehouse. Each app connection includes: * **Service account details** - Username and authentication method * **Database configuration** - Target database, schema, and warehouse * **Role and permissions** - What the connection can access and modify * **Connection parameters** - Timeout settings, SSL configuration, etc. For example, your ingestion capability needs an app connection with write permissions to load raw data, while your BI capability needs a separate app connection with read-only access to query transformed data. ## What are credentials? Credentials are your personal authentication details for accessing the data warehouse through the 5X IDE. Unlike app connections which are shared across your workspace for automated processes, credentials are individual to you and enable interactive development work. Your credentials allow you to: * Execute SQL queries in the 5X IDE * Run and test dbt models during development * Browse database schemas and tables * Perform ad-hoc data analysis ## Do you need to configure app connections? App connections are essential for all 5X workspaces, but your setup approach depends on how you provisioned your workspace: ### If you're using a 5X managed warehouse **You can start with defaults and customize later.** When you provision a workspace with a 5X managed warehouse, the platform automatically creates default app connections for all core capabilities based on recommended best practices. These defaults allow you to start using 5X immediately. **However, you'll likely want to re-configure/modify app connections as you grow:** * **Connect to different databases or schemas** for specific projects or environments * **Change roles or permissions** to match your organization's security requirements * **Create multiple connections** when working with multiple databases or data sources * **Set up environment-specific connections** for development, staging, and production workflows * **Customize connection parameters** like warehouses, or authentication methods **For getting started:** You can proceed with the default setup and return to configure additional app connections as your needs evolve. ### If you're using an existing warehouse **You must configure all app connections manually.** When you connect an existing Snowflake, BigQuery, AWS Redshift, or PostgreSQL warehouse, no default app connections are created. This makes app connection configuration critical for basic platform functionality. **What you need to set up:** * **Before using any 5X capability**, you need to configure its corresponding app connection * **Each capability requires its own connection** with appropriate permissions and settings * **Without proper app connections**, capabilities like data ingestion, dbt modeling, and BI dashboards won't function **What this means practically:** * Want to ingest data? You need an app connection of type **"Ingestion"** * Want to run dbt models or orchestrate workflows? You need an app connection of type **"Job"** * Want to create dashboards? You need an app connection of type **"Business Intelligence"** * Want to use metrics? You need an app connection of type **"Metrics Store"** * Want to use Streamlit-based data apps? You need an app connection of type **"Data & AI Apps"** This is why understanding app connections is critical when using an existing warehouse: they're the foundation that enables all 5X platform functionality. ## Creating app connections ### 1. Access app connections 1. **Navigate to settings** * From your workspace, click **"Settings"** in the left sidebar * Select **"App connections"** from the settings menu Settings page with app connections highlighted in sidebar 2. **Create new connection** * Click **"+ New connection"** * You'll see the connection configuration form ### 2. Configure connection details The configuration fields vary depending on your warehouse type and the capability you're setting up. Here's how to configure each field: **Connection types and fields:** **For all Snowflake connections, you'll configure:** 1. **Type:** Select the capability type * **Ingestion** - For data ingestion pipelines * **Job** - For dbt modeling and orchestration workflows * **Business Intelligence** - For BI dashboards and analytics * **Metrics Store** - For metrics layer operations * **Data & AI Apps** - For data and AI applications 2. **Connection name:** Use descriptive names like: * `prod_ingestion` * `dev_modeling` * `bi_dashboard` * `metrics_store` * `data_apps` 3. **Snowflake account:** Your Snowflake account URL * Format: `https://[account].snowflakecomputing.com` 4. **Default role:** The role this connection will use * Examples: `INGESTION_ROLE`, `DBT_ROLE`, `BI_ROLE` 5. **Default database:** Target database name * Examples: `ANALYTICS_DB`, `PROD_DB`, `DEV_DB` 6. **Default schema:** Default schema for operations * Examples: `RAW_DATA`, `TRANSFORMED`, `MARTS` 7. **Default warehouse:** Compute warehouse to use * Examples: `INGESTION_WH`, `TRANSFORM_WH`, `BI_WH` **Authentication options:** * **Key pair authentication:** Service account username with private key (recommended for production) **Pre-requisites and setup:** For detailed setup instructions including service account creation, role configuration, and permissions, refer to the "Pre-requisites" section on the app connection setup page as shown in the screenshot. Snowflake Business Intelligence app connection configuration showing all required fields **Connection types and fields:** **For all BigQuery connections, you'll configure:** 1. **Type:** Select the capability type * **Ingestion** - For data ingestion pipelines * **Job** - For dbt modeling and orchestration workflows * **Business Intelligence** - For BI dashboards and analytics * **Metrics Store** - For metrics layer operations * **Data & AI Apps** - For data and AI applications 2. **Connection name:** Use descriptive names like: * `prod_ingestion` * `dev_jobs` * `bi_analytics` * `data_apps` 3. **Upload file:** Service account JSON key file * Download from Google Cloud Console > IAM & Admin > Service Accounts * Must have appropriate BigQuery permissions 4. **Project ID:** Your Google Cloud project identifier 5. **Dataset location:** Geographic location of your datasets * Examples: `US`, `EU`, `us-central1` **For Job connections (dbt/orchestration):** * **Dataset name:** Target dataset for dbt models * **dbt version:** Select your preferred dbt version **Pre-requisites and setup:** For detailed setup instructions including service account creation, role assignment, and JSON file generation, refer to the "Pre-requisites" section on the app connection setup page as shown in the screenshots. BigQuery ingestion app connection configuration showing service account upload and project settings **Key considerations for all warehouse types:** * **Use service accounts** instead of personal credentials for app connections * **Follow least privilege principle** - grant only necessary permissions for each capability * **Use descriptive naming** to easily identify connection purposes * **Test connections** before saving to ensure proper configuration ### 3. Connect and create app connection 1. **Click Connect** * After configuring all required fields, click the **"Connect"** button * This will first test your connection configuration automatically * If the test is successful, the app connection will be created and saved 2. **Connection validation** * The system verifies authentication credentials and database access * Checks that the specified role has necessary permissions * Ensures all connection parameters are valid 3. **Success confirmation** * Once created successfully, the connection will appear in your app connections list * You can now use this connection for the selected capability ## Managing user credentials ### 1. Access credentials settings 1. **Navigate to credentials** * Go to **Settings > Credentials** * Click **"+ Add credential"** or edit existing credentials Credentials page showing Snowflake authentication options ### 2. Configure personal credentials Based on your warehouse type, configure your personal access: **Required fields:** 1. **Snowflake Account:** Your Snowflake account URL * Format: `https://nnb42337.us-east-1.snowflakecomputing.com` 2. **Default Role:** Your personal role for queries * Examples: `ANALYST`, `DEVELOPER`, `SYSADMIN` 3. **Default Warehouse:** Warehouse for your interactive queries * Examples: `COMPUTE_WH`, `DEV_WH`, `ANALYST_WH` 4. **Authentication Type:** Choose your authentication method * **Username & Password:** Standard authentication with your Snowflake credentials * **Key Pair:** More secure option using private key authentication **dbt Integration (optional):** If you plan to use dbt modeling: 1. **Enable dbt toggle:** Turn on the "dbt" toggle in credentials * This enables dbt operations in the 5X IDE 2. **dbt version:** Select your preferred dbt version * Available versions: `1.6.18`, `1.7.19`, `1.8.9`, `1.9.10` 3. **Database:** Target database for dbt models 4. **Schema:** Schema for dbt transformations 5. **(optional field) Target name:** Environment identifier for dbt * Examples: `dev`, `prod`, `staging` **Required fields:** 1. **Authentication:** Sign in with Google * Click **"Sign in with Google"** to authenticate with your Google account * Select the Google account that has access to your BigQuery project * Grant necessary permissions for 5X to access BigQuery on your behalf 2. **Project ID:** Your Google Cloud project identifier * Example: `cmp-princestackprovision-czwf` **dbt Integration (optional):** If you plan to use dbt modeling: 1. **Enable dbt toggle:** Turn on the "dbt" toggle in credentials * This enables dbt operations in the 5X IDE 2. **dbt version:** Select your preferred dbt version * Available versions: `1.6.18`, `1.7.19`, `1.8.9`, `1.9.10` 3. **Database:** Target database for dbt models 4. **Schema:** Schema for dbt transformations 5. **(optional field) Target name:** Environment identifier for dbt * Examples: `dev`, `prod`, `staging` ## Best practices ### Naming conventions **Recommended naming:** * `[Environment]_[Capability]_Connection` * Examples: * `Prod_Ingestion_Connection` * `Dev_Modeling_Connection` * `Staging_BI_Connection` **Recommended naming:** * `[CAPABILITY]_[ENVIRONMENT]_USER` * Examples: * `INGESTION_PROD_USER` * `DBT_DEV_USER` * `BI_STAGING_USER` ### Security guidelines **Security best practices:** * **Use least privilege principle** - Grant only necessary permissions for each capability * **Separate environments** - Use different connections for dev/staging/prod * **Rotate credentials regularly** - Implement credential rotation policies * **Monitor connection usage** - Review access logs and connection patterns ## Troubleshooting common issues **Common causes:** * Incorrect connection details * Invalid credentials or expired passwords * Network connectivity issues * Insufficient permissions for the specified role **Solutions:** * Verify the configured credentials * Check role permissions in your warehouse * Ensure network access and correct firewall settings * Test credentials directly in your warehouse console **Common causes:** * Role lacks necessary database permissions * Warehouse access restrictions * Schema or table-level permission issues **Solutions:** * Review and grant necessary permissions to the role * Check warehouse access policies * Verify schema and table permissions * Contact your database administrator for access issues **Common causes:** * Incorrect private key format * Missing or incorrect passphrase * Expired or revoked credentials **Solutions:** * Verify private key format (PEM format required) * Check passphrase accuracy * Regenerate keys if expired * Ensure proper key pair setup in warehouse ## What's next? With your app connections and credentials configured, you're ready to start using the 5X Platform's capabilities. **Next:** Ingest your data Configure your first data ingestion pipeline using the app connections you just set up. *** Return to quickstart guide overview Steps 1-3: Account setup and workspace provisioning # Step 5: Ingest Your Data Source: https://documentation.5x.co/quickstart/step-5-ingest-data Connect and ingest data from 600+ sources into your 5X platform with seamless integration capabilities **Estimated Time:** 15-30 minutes Configure your first data ingestion pipeline and start bringing data from your various sources into the 5X platform. ## Ready to ingest data With your app connections now configured, you're now ready to set up your first data ingestion pipeline and start bringing data into your warehouse. ## What is data ingestion? Data ingestion is the process of bringing raw data from external systems into your warehouse, standardizing its format and centralizing it for further use. Whether you're pulling sales data from a CRM, logs from internal databases, or campaign performance from a marketing platform, 5X enables you to ingest with minimal engineering effort. The 5X platform provides a managed, secure, and no-code user experience for data ingestion, eliminating the complexity of managing multiple data integration tools. ## Your ingestion toolkit: 600+ connectors 5X supports a broad catalog of 600+ out-of-the-box connectors, including: **Popular SaaS applications:** * Salesforce, HubSpot, NetSuite * Zendesk, Shopify, Stripe * Slack, Jira, Confluence **All major database systems:** * PostgreSQL, MySQL, MongoDB * Redshift, Snowflake, BigQuery * Oracle, SQL Server, Cassandra **Analytics and marketing platforms:** * Google Analytics, Facebook Ads * Mixpanel, Klaviyo, Mailchimp * Adobe Analytics, LinkedIn Ads **File and cloud storage:** * Google Sheets, Excel (via cloud) * Amazon S3, Google Cloud Storage * FTP, SFTP, Dropbox If a connector is not available in our catalog, you can request custom connectors, which are typically delivered within days. You can also ingest from modern APIs or legacy sources using custom integration blueprints maintained by 5X. ## Setting up your first ingestion pipeline Let's walk through creating your first data ingestion pipeline: ### 1. Access the ingestion module 1. **Navigate to Ingestion** * Visit [platform.5x.co](https://platform.5x.co) and log into your workspace * In the left sidebar, click on **Ingestion** 5X Ingestion Dashboard 2. **Start New Connector Setup** * On the ingestion dashboard, click **Add Connector** to begin setup ### 2. Select your data source 1. **Choose from 600+ Connectors** * Use the search bar to find your desired source (e.g., "Google Sheets" or "Salesforce") * Click the source card to begin the configuration process Connector Selection Interface 2. **Select a Destination** Connector Selection Interface 3. **Enter Destination Schema** Connector Selection Interface ### 3. Configure your connector The setup screen will prompt you for several fields. These vary by connector, but common fields include: Connector Configuration Form **Required Information:** * **Destination schema name:** Choose where data will be written in your warehouse * **Connection parameters:** Configure source-specific connection settings * **Data selection:** Choose which objects, tables, or datasets to sync * **Additional settings:** Configure any source-specific parameters as required **Note:** The connector name is automatically generated based on your destination schema name. **Authentication varies by connector type and may include:** * **OAuth flows:** Direct login with your account credentials * **API keys and tokens:** Service account credentials or API authentication * **Database credentials:** Username, password, and connection strings * **Certificate-based:** Private keys or certificate files for secure connections * **Resource-specific access:** Direct URLs or specific resource identifiers The system will guide you through the appropriate authentication method for your selected source. **Additional configuration options may include:** * **Data filtering and transformation:** Apply filters or basic transformations during ingestion * **Security settings:** Configure data privacy and access controls * **Performance tuning:** Optimize sync behavior for your specific use case * **Compliance features:** Enable data governance and audit capabilities Available options depend on the specific connector and source system capabilities. Use descriptive destination schema names since they determine your connector names and where data lands in your warehouse. Good examples: `salesforce_crm`, `google_analytics`, `postgres_customers`, `shopify_orders`. Avoid spaces and use underscores for multi-word schemas. ### 4. Complete setup and start initial sync 1. **Review Configuration** * Double-check all settings and authentication details * Verify data selection and sync frequency * Preview the sync configuration 2. **Trigger Initial Sync** * Click **Continue** through each configuration screen * Once setup is complete, you'll see a confirmation screen * Click **Start Initial Sync** to begin your first data transfer * The initial sync will run immediately to populate your warehouse with data Connector Activation Screen ### 5. Manage sync settings and data selection Once your connector is active, you can manage its sync behavior and data selection from the connector details page: **Access Sync Settings:** * From the ingestion dashboard, click on your connector name * Or use the Actions menu to access connector settings * Manual Sync: Trigger syncs on-demand using the "Sync now" button **Sync Frequency Options:** Configure how often your data syncs using either: * **Fixed Intervals:** Choose from preset options (1 minute to 24 hours) * **CRON Expressions:** Set custom schedules using CRON syntax **Schema and Table Selection:** * Use the **Schema** tab to manage which data gets synced * Select or deselect entire tables with checkboxes * Choose specific columns within each table * Use search functionality to find specific tables quickly * Toggle "Show selected tables" to focus on active data sources **Connector Management:** * **Test connection:** Verify your source connection is working * **Sync now:** Trigger an immediate sync outside the schedule * **Pause:** Temporarily stop all syncing ### Best practices for secure ingestion **Best Practices:** * Use service accounts instead of personal credentials when possible * Rotate API keys and credentials regularly * Apply principle of least privilege for data access * Monitor credential usage and access patterns **Governance Recommendations:** * Hash or exclude sensitive fields (PII, financial data) during ingestion * Set up data retention policies for ingested data * Document data sources and their business purposes * Regular review and audit of active connectors **Performance Tips:** * Choose appropriate sync frequencies based on data change rates * Use incremental sync for large datasets * Schedule resource-intensive syncs during off-peak hours ## Troubleshooting common issues **Common Issues:** * Network connectivity problems * Authentication failures (expired tokens, wrong credentials) * Source system downtime or maintenance * Firewall or security group restrictions **Solutions:** * **Verify connection settings:** Double-check host, port, and database names * **Test authentication:** Ensure credentials are valid and have proper permissions * **Check network connectivity:** Verify firewall rules and network access * **Monitor source system status:** Check if the source system is operational * **Review error logs:** Look for specific error messages in sync history **Common Problems:** * Data type mismatches between source and destination * Unexpected null values in required fields * Character encoding issues (special characters, unicode) * Date format inconsistencies across systems * Duplicate records or primary key violations **Solutions:** * **Review data type mappings:** Ensure compatible data types * **Implement data validation:** Set up rules to catch quality issues * **Add data cleaning:** Use transformations to standardize data * **Monitor data quality:** Set up alerts for quality degradation * **Document data quirks:** Note known issues and workarounds ## What's next? With ingestion complete and your data flowing into the warehouse, you can now: Query and explore your ingested data using the integrated SQL editor Add teammates and set up permissions for collaborative data work Discover advanced features when you're ready to go deeper Ingestion is the critical first step in the 5X data lifecycle. With proper setup, it enables trusted, timely, and scalable data access across your organization. **Next:** Explore your data Query and explore your ingested data using the integrated SQL editor to understand its structure and quality. *** Return to quickstart guide overview Step 4: App connections and credentials # Step 6: Explore Your Data Source: https://documentation.5x.co/quickstart/step-6-explore-data Query and explore your ingested data using the integrated SQL editor and schema browser **Estimated Time:** 10-15 minutes Use the integrated SQL editor to explore your ingested data and validate that everything was imported correctly. ## Accessing the IDE Now that your data has been ingested, it's time to explore it using the 5X platform's modern VS Code-based IDE: 1. **Navigate to IDE** * Click **"IDE"** in the left sidebar of your workspace * Click **"Start Session"** to initialize your development environment * It will take sometime for the environment to fully load 5X IDE Start Session Interface **For existing warehouse users:** If you connected your own existing warehouse (Snowflake, BigQuery, AWS Redshift, or PostgreSQL) instead of using the 5X managed option, you'll need to configure your IDE credentials in **Settings → IDE → Credentials** before you can query your data. The IDE will guide you through this setup process. ## Exploring your data Once you have access to the IDE, you can start exploring your ingested data: ### 1. Browse your data sources The **Database Explorer** in the IDE sidebar provides direct access to your data warehouse: * **Browse schemas**: Expand database nodes to see available schemas and datasets * **Explore tables**: Click on table names to view column structures, data types, and metadata * **Preview data**: Use the preview functionality to see sample data from your tables * **Execute queries**: Right-click on tables to generate SELECT statements IDE Database Explorer with Schema Navigation ### 2. Write your first queries Use the SQL editor to run queries against your data. Here are some useful starter queries: **Quick data preview:** ```sql theme={null} -- View the first 10 rows from any table SELECT * FROM .. LIMIT 10; ``` **Check data volume:** ```sql theme={null} -- See how much data was ingested SELECT COUNT(*) as total_rows FROM ..; ``` ### 3. Using the VS Code SQL editor The IDE provides a full-featured VS Code experience with specialized SQL development capabilities: * **SQL editor**: Write and execute SQL queries with syntax highlighting and IntelliSense * **Results panel**: View query results in an interactive table format with pagination * **Multiple tabs**: Work with multiple queries simultaneously using VS Code tabs * **Auto-completion**: Get intelligent suggestions for table names, columns, and SQL keywords * **Query execution**: Execute queries directly from the editor with keyboard shortcuts * **Export functionality**: Download query results in CSV, JSON, or Excel formats * **Query history**: Access previously executed queries for reuse **Quick tips:** * Use `Ctrl+Enter` (Windows/Linux) or `Cmd+Enter` (Mac) to execute queries * Right-click on tables in the Database Explorer to generate SELECT statements * Use the integrated terminal for advanced SQL operations and dbt commands ## Troubleshooting common issues **Possible causes:** * Browser compatibility issues * Network connectivity problems * Platform maintenance * Selected compute profile is below the minimum requirements **Solutions:** * Try refreshing your browser and starting a new session * Check your internet connection * Use a supported browser (Chrome, Firefox, Safari, Edge) * Contact support if issues persist **Possible causes:** * Data ingestion may still be in progress * IDE credentials not configured (for existing warehouse users) * Schema permissions issues **Solutions:** * Check ingestion status in the Ingestion page * Configure IDE credentials in Settings → IDE → Credentials * Verify your warehouse permissions * Try restarting your IDE session **Common issues:** * Syntax errors in SQL queries * Missing permissions for specific tables * Warehouse connection timeout **Solutions:** * Review query syntax carefully * Check error messages for specific guidance * Use the Database Explorer to verify table names and schemas * Try restarting your IDE session if connections seem stale **Common issues:** * Slow response times * High memory usage * Session timeouts **Solutions:** * Close unused files and tabs * Restart your IDE session using the controls in the top-left * Limit large query results using LIMIT clauses * Clear browser cache if issues persist ## What's next? Great! You've successfully explored your ingested data. Now you're ready to transform it into analysis-ready datasets: Add teammates and set up permissions for collaborative data work Discover advanced features when you're ready to go deeper The IDE is your primary tool for working with data on the 5X platform. Take some time to familiarize yourself with the interface and explore your data before inviting your team to collaborate. **Next:** Invite your team Add teammates to your workspace and set up collaboration for your data projects. *** Return to quickstart guide overview Step 5: Ingest your data # Step 7: Invite Your Team Source: https://documentation.5x.co/quickstart/step-7-invite-team Add teammates to your workspace and set up basic permissions for collaboration **Estimated Time:** 5 minutes Invite your teammates and set up basic permissions to enable collaborative data work on your 5X workspace. ## Why invite your team? Data projects are most successful when teams can collaborate effectively. By inviting your teammates early, you'll: * **Enable collaborative data exploration** - Multiple team members can query and analyze data together * **Share insights faster** - Everyone can access the same data sources and results * **Establish proper governance** - Set appropriate access levels from the start * **Accelerate adoption** - More users means faster time to value across your organization Team Collaboration in 5X Workspace ## Adding team members ### 1. Access team management 1. **Navigate to Team management section** * Click **"Team"** in the left sidebar * You'll see the team management interface with **Users** and **Roles** tabs 2. **Team Management Interface** * **Users tab**: View current workspace members and their roles * **Roles tab**: Manage role permissions (if needed) * See both 5X platform roles and warehouse-specific roles * Track user status and manage access ### 2. Invite new users 1. **Add Individual Users** * Click **"Add new user"** button in the top right * Fill in the user's **Name** and **Email** * Select their **5X Role** (Admin, Contributor, etc.) * Choose **Warehouse Roles** based on their data access needs * Click **"Add new user"** to send the invitation Add New User Dialog ## Understanding user roles The 5X platform uses two types of roles: **5X Platform Roles** - Control access to platform features: * **Admin** - Full platform management and user control * **Developer** - Build models, queries, and data applications * **Member** - Standard access for data work and collaboration * **BI User** - Focus on dashboards and business intelligence * **Custom** - Define specific permission sets **Warehouse Roles** - Control access to your data warehouse for querying and data operations. **Keep it simple:** Start with basic role assignments - you can always adjust permissions later as your team's needs evolve. **Tip:** New users will receive an email invitation to join your workspace. They'll need to accept the invitation to gain access to the platform. ## The invitation process ### What happens when you invite someone 1. **Invitation Email Sent** * Users receive an invitation email from "The 5X Team" * Email includes your workspace name and a brief platform description * Contains an "Accept Invite" button for easy access 2. **User Experience** * Recipients click "Accept Invite" to join your workspace * They'll be guided through any necessary account setup * Once accepted, they'll have access based on the roles you assigned 5X Invitation Email ## What's next? Congratulations! You've successfully set up your 5X workspace with data ingestion, exploration capabilities, and team collaboration. Your team now has everything they need to start getting value from the platform. Learn about advanced features you can explore when ready Find answers to common questions and get help ## You're ready to go! Your 5X workspace is now fully operational with: * ✅ **Data connected and ingesting** from your sources * ✅ **Team members invited** with appropriate permissions * ✅ **SQL IDE ready** for data exploration and analysis * ✅ **Foundation set** for collaborative data work Your team can now start querying data, sharing insights, and collaborating effectively. The platform will continue to ingest data from your connected sources, keeping everything up to date automatically. **Optional:** Explore advanced features and capabilities When you're ready to go deeper, learn about data modeling, dashboards, orchestration, and more. *** Return to quickstart guide overview Step 6: Explore your data # Step 8: What's Next? Source: https://documentation.5x.co/quickstart/step-8-next-steps Explore advanced features and capabilities when you're ready to go deeper **Congratulations!** 🎉 You've successfully completed the essential 5X setup. Your workspace is now operational with data ingestion, SQL exploration, and team collaboration ready to go. This section is **optional** - explore these advanced capabilities when you're ready to unlock more power from the platform. ## You've completed the essentials Your 5X workspace now has everything you need to start getting value: ✅ **Account and workspace** fully configured\ ✅ **Data sources connected** and actively ingesting\ ✅ **SQL IDE ready** for data exploration and analysis\ ✅ **Team members invited** with appropriate permissions **You can stop here and start working with your data, or continue exploring advanced features below.** ## Advanced capabilities to explore When you're ready to go deeper, these powerful features are available: ### Data modeling and transformation Transform your raw data into analysis-ready datasets: **What it enables:** * Transform raw data into structured models * Create reusable data transformations * Generate automated documentation * Track data lineage and dependencies **Best for:** Creating clean, reliable datasets for analysis **What it enables:** * Write custom SQL transformations * Create views and materialized tables * Build complex data pipelines * Schedule data processing jobs **Best for:** Custom data preparation workflows ### Visualization and dashboards Build interactive dashboards and data applications: **What it enables:** * Drag-and-drop dashboard creation * 50+ chart types and visualizations * Self-service analytics for teams * Role-based dashboard sharing **Best for:** Executive dashboards and standard reporting **What it enables:** * Custom data applications * Interactive tools and calculators * Machine learning model demos * Specialized workflows **Best for:** Custom tools and specialized use cases ### Workflow automation Automate your data pipelines and processes: * Visual pipeline creation * Automated job scheduling * Data quality monitoring * Dependency management **Best for:** Complex data workflows and automation ## Ready to dive deeper? Your 5X workspace is ready for whatever comes next. Whether you start with advanced features today or return to them later, the foundation is solid. Return to the quickstart overview or continue exploring your workspace *** Step 7: Invite your team Explore the full platform documentation