kmeans-seeding

User Guide

  • Installation
    • Basic Installation
    • Requirements
    • Installing with FAISS Support
    • Installing from Source
    • Platform-Specific Notes
      • macOS
      • Linux
      • Windows
    • Verifying Installation
    • Troubleshooting
      • C++ Extension Not Found
      • FAISS Not Found
      • OpenMP Warnings
    • Upgrading
    • Uninstalling
  • Quickstart Guide
    • Basic Usage
    • With Scikit-Learn
    • Trying Different Algorithms
      • Standard k-means++
      • RS-k-means++ (Fast & Accurate)
      • AFK-MC² (MCMC Sampling)
      • Fast-LSH k-means++ (Google 2020)
    • Reproducibility
    • Complete Example
    • Comparing Algorithms
    • Next Steps
  • Choosing an Algorithm
    • Start Here: Quick Selection
    • By Dataset Size
      • Small (n < 10,000)
      • Medium (10,000 < n < 100,000)
      • Large (n > 100,000)
    • By Dimensionality
      • Low (d < 20)
      • Medium (20 < d < 100)
      • High (d > 100)
    • By Number of Clusters
      • Few (k < 50)
      • Medium (50 < k < 500)
      • Many (k > 500)
    • By Data Type
      • Dense Numerical
      • Sparse (Text, One-Hot)
      • Images/Embeddings
      • Time Series
    • By Requirements
      • Maximum Speed
      • Best Quality
      • No Dependencies (No FAISS)
      • Simple Setup
    • Summary Flowchart
    • See Also
  • Scikit-Learn Integration
    • Basic Integration
      • Key Points
    • Multiple Initializations
    • Pipeline Integration
    • With MiniBatchKMeans
    • Cross-Validation
    • Feature Preprocessing
    • Evaluation Metrics
    • See Also

Algorithm Documentation

  • RS-k-means++: Rejection Sampling
    • Overview
    • Algorithm Details
    • Python API
      • rskmeans()
    • Index Type Comparison
    • Parameter Tuning
      • max_iter
      • index_type
    • Performance Tips
    • Theoretical Background
    • References
    • See Also
  • AFK-MC²: Adaptive Fast k-MC²
    • Overview
    • Algorithm Details
    • Python API
      • afkmc2()
    • How It Works
      • Markov Chain Construction
      • Sampling Process
    • Parameter Tuning
      • chain_length
      • Adaptive Chain Length
    • Performance Characteristics
      • Time Complexity
      • Quality vs Speed Tradeoff
    • Comparison with Other Algorithms
      • vs Standard k-means++
      • vs RS-k-means++
    • Practical Tips
    • Theoretical Background
    • References
    • See Also
  • Fast-LSH k-means++
    • Overview
    • Algorithm Details
    • Python API
      • multitree_lsh()
    • Parameter Tuning
      • n_trees
      • n_greedy_samples
      • scaling_factor
    • How It Works
      • Tree Embedding
      • Sampling Process
      • Integer Casting
    • Performance Characteristics
      • When Fast-LSH Excels
    • Comparison with Other Algorithms
    • Practical Tips
    • Theoretical Background
    • References
    • See Also
  • Standard k-means++
    • Overview
    • Python API
      • kmeanspp()
    • Algorithm
      • D² Sampling
    • Advantages
    • Limitations
    • When to Use Alternatives
    • See Also
  • Algorithm Comparison
    • Quick Recommendation
    • Decision Tree
    • Detailed Comparison
      • Feature Matrix
      • Performance Benchmarks
    • Use Case Recommendations
      • Text Clustering (TF-IDF, Word2Vec)
      • Image Clustering (Features/Embeddings)
      • Customer Segmentation
      • Time Series / Sensor Data
      • Biological / Genomic Data
      • Small Datasets (n < 10K)
    • Quality vs Speed Tradeoff
      • Fastest (Slight Quality Loss)
      • Balanced (Recommended)
      • Best Quality (Slower)
      • Exact (Slowest)
    • Common Patterns
      • Pattern 1: Try Fast First
      • Pattern 2: Multiple Runs
      • Pattern 3: Progressive Refinement
    • FAQ
    • Summary Table
    • See Also

API Reference

  • API Reference
    • Main Functions
      • rskmeans
        • rskmeans()
      • kmeanspp
        • kmeanspp()
      • afkmc2
        • afkmc2()
      • multitree_lsh
        • multitree_lsh()
    • Aliases
      • rejection_sampling
      • fast_lsh
    • Complete Example
    • Parameter Summary
      • Common Parameters
      • Algorithm-Specific Parameters
    • Return Values
    • Exceptions
    • Type Hints
    • See Also
  • Advanced Usage
    • Custom Distance Metrics
    • Incremental Clustering
    • Parallel Initialization
    • Hierarchical Initialization
    • Weighted Sampling
    • Warm Starting
    • Memory-Efficient Processing
    • Custom Index Parameters (RS-k-means++)
    • Debugging and Profiling
    • See Also

Additional Information

  • Changelog
    • Version 0.2.2 (November 2025)
    • Version 0.2.1 (2025)
    • Version 0.2.0 (2025)
    • Version 0.1.0 (2024)
    • Future Plans
  • Contributing
    • Getting Started
    • Development Setup
    • Code Style
    • Testing
    • Documentation
    • Pull Request Process
    • Reporting Issues
    • Feature Requests
    • Code of Conduct
    • License
    • Contact
    • Thank You!
  • References
    • Academic Papers
      • RS-k-means++
      • k-means++
      • AFK-MC²
      • Fast-LSH k-means++
    • Related Work
      • Clustering
      • Approximate Nearest Neighbors
    • External Resources
      • Software
      • Documentation
      • Tutorials
    • Citation
    • Contact
kmeans-seeding
  • Search


© Copyright 2025, Poojan Shah.

Built with Sphinx using a theme provided by Read the Docs.