I've been trying to contribute to statrs recently, and one of the issues I've been running into is that I'm trying to have an iterator that returns items in sorted order without cloning the underlying data. I'm working this PR if you wanted to take a look at my code so far. I looked up this problem and found this StackOverflow post about this topic from years ago but, frankly, I don't believe it. Assuming you're iterating from a vector, mutating the underlying vector is akin to keeping state on the order of the vector. Why cant that happen in a separate data structure? Is there a more efficient way to represent ordinality other than a vector? Creating an iterator is simply tracking traversal through that structure, which I think could be done using a bloom filter to track which indices have not been traversed yet. That just leaves the traversal algorithm itself. What information would an iterator need to know to make the best decision? Could I adapt a sorting algorithm to be an iterator?
I'm asking a lot of questions because I'm a statistics guys, not an algorithms guy. Any starting point or input is much appreciated!
If you have a backing storage and random access, you can implement any sorting function and simply yield/return the values in order rather than commit them into a collection. For example, you can track the minimum value after one full iteration, then iterate in reverse tracking both the current minimum and the next minimum and yielding each current minimum, then repeat going the other direction, and so on until you're returning only the maximums. In theory, this shouldn't require any internal buffer at all, but does require a total ordering constraint on the type (and is
O(n^2)and in practice slow af)