Functional limit theorems for the number of occupied boxes in the Bernoulli sieve Gerold Alsmeyer1, Alexander Iksanov2 and Alexander Marynych1,2 6 1 0 2 n a J 7 Abstract TheBernoulli sieveis theinfiniteKarlin “balls-in-boxes”scheme with random probabilities 1 of stick-breaking type. Assuming that the number of placed balls equals n, we prove several functional ] lniummitbetrheKor∗e(mt)so(fFbLoTxse)sicnontthaeinSinkogroahtomdosspta[cnet]Db[a0l,ls1,]ten∈d[o0w,e1d], wanitdhththeeraJn1d-oomr Mdis1t-rtiobpuotlioognyfufonrcttiohne n R K∗(t)/K∗(1), as n → ∞. The limit processes for K∗(t) are of the form (X(1)−X((1−t)−)) , n n n t∈[0,1] P where X is either a Brownian motion, a spectrally negative stable L´evy process, or an inverse stable . subordinator. The small values probabilities for the stick-breaking factor determine which of the alter- h natives occurs. If the logarithm of this factor is integrable, the limitprocess for K∗(t)/K∗(1) is a L´evy n n at bridge. Our approach relies upon two novel ingredients and particularly enables us to dispense with a m Poissonization-de-Poissonization step which hasbeen an essential component in all theprevious studies of K∗(1). First, for any Karlin occupancy scheme with deterministic probabilities (p ) , we obtain n k k≥1 [ an approximation, uniformly in t∈[0,1], of the number of boxes with at most [nt] balls by a counting function defined in terms of (p ) . Second, we prove several FLTs for the number of visits to the in- 1 k k≥1 terval[0,nt]byaperturbedrandomwalk,asn→∞.Ifthestick-breakingfactorhasabetadistribution v with parameters θ >0 and 1, the process (K∗(t)) has the same distribution as a similar process 4 n t∈[0,1] definedbythenumberofcyclesof lengthatmost[nt]inaθ-biasedrandompermutation a.k.a.aEwens 7 permutation with parameter θ. As a consequence, our FLT with Brownian limit forms a generalization 2 of a FLT obtained earlier in the context of Ewens permutations by DeLaurentis and Pittel [5], Hansen 4 [17], Donnelly,Kurtz andTavar´e [6], and Arratia and Tavar´e [3]. 0 . 1 AMS 2000 subject classifications: primary 60F17; secondary 60C05 0 Keywords: Bernoulli sieve, infinite urn model, perturbed random walk, renewal theory 6 1 : v i 1 Introduction X r a 1.1 The Bernoulli sieve and regenerative compositions Given a sequence W ,W ,... of independent copies of a (0,1)-valued random variable W, 1 2 consider the random partition of (0,1] into the subintervals (V ,V ], i N, called boxes i i−1 ∈ hereafter, where n V := 1, and V := W for n N. 0 n i ∈ i=1 Y 1 InstituteofMathematicalStatistics,DepartmentofMathematicsandComputerScience,Universityof Mu¨nster, Orl´eans-Ring 10, D-48149 Mu¨nster, Germany. 2 Faculty of Cybernetics, Taras ShevchenkoNational Universityof Kyiv,01601 Kyiv,Ukraine e-mail: [email protected],[email protected],[email protected] 1 2 Gerold Alsmeyer1, Alexander Iksanov2 andAlexander Marynych1,2 One may also think of the V as the successive cut points in a stick-breaking procedure at i the end of which a stick with endpoints 0 and 1 is broken into the infinitely many pieces with endpoints Vi and Vi−1, i N. Next, let (Uj)j∈N be a sequence of independent uniform ∈ (0,1)randomvariableswhichisalsoindependentof(Wk)k∈N.TheBernoulli sieve isarandom occupancy scheme in which balls, labeled by 1,2,... and of random weights U ,U ,..., are 1 2 placed into the boxes in accordance with their weight. Hence, a ball of weight U goes into the box (V ,V ] iff V < U V . Since its introduction by Gnedin [7], the model has been i i−1 i i−1 ≤ studied in a series of articles [9, 10, 11, 12, 13, 18, 19, 22]. The term “Bernoulli sieve” can be understood when interpreting the occupancy scheme in terms of a randomized leader-election procedure (see [7] for more details). Defining the (infinite) vector of occupancy counts Z∗ := # 1 j n:U (V , V ] , i N, (1) n,i { ≤ ≤ j ∈ i i−1 } ∈ thus Zn∗,i ≥ 0 for i ∈ N and i≥1Zn∗,i = n, we see that the Zn∗ := (Zn∗,i)i∈N induces a weak composition of the integer n. The attribute “weak” is intended to emphasize that zero parts P of the composition are allowed. By sequentially allocating the points U ,U ,..., the sequence 1 2 ( ∗) isconsistentlydefinedforalln N ,anditfurtherhasthefollowingtwodistinguished Zn n≥0 ∈ 0 regenerative properties: Sampling consistency: If one out of n points is removed uniformly at random from the • interval it belongs to, then the resulting weak composition of n 1 has the same law as − ∗ . Zn−1 The deletion property: If the first interval (V ,1] contains m points and is removed, then a 1 • weak composition of n m with the same law as ∗ is obtained. − Zn−m The classofrandomcompositionsgeneratedby aBernoullisievedoesnotcoverallregener- ativecompositions(i.e.,thosehavingthetwoaforementionedproperties).Infact,Theorem5.2 in [14] states that every consistent family of regenerative compositions can be constructed by allocating the points U ,U ,... to countably many open intervals forming the complement of 1 2 the closed range of a multiplicative zero-drift subordinator (e−Lt)t≥0 independent of the uni- formsample.Withinthismoregeneralframework,weakcompositionspertainingtoaBernoulli sieve are those corresponding to a compound Poissonprocess (L ) . t t≥0 IntheclassicaloccupancyschemeofKarlin[23]ballsareplacedindependentlyinaninfinite arrayofboxesinaccordancewithaprobabilityvector(pk)k∈N,wherepkdenotestheprobability ofchoosingboxk.TheBernoullisieveistheKarlinoccupancyschemewithrandomprobability p∗ = V V = W W (1 W ) (2) k k−1− k 1··· k−1 − k of choosing box k ∈ N and such that, given (p∗k)k∈N, balls are allocated independently. The Bernoulli sieve can therefore be thought of as an occupancy scheme in random environment, the latter being defined by the i.i.d. random variables W ,W ,.... 1 2 For r=1,...,n, denote by Kn∗,r := i≥11{Zn∗,i=r} the number of boxes containing exactly r balls, and let Kn∗ := nr=1Kn∗,r = Pi≥11{Zn∗,i≥1} denote the number of nonempty boxes. Then define P P[nt] Kn∗(t) := Kn∗,r = 1{Z∗ ∈[1,nt]} n,i r=1 i≥1 X X for t [0,1]. ∈ Throughout the rest of the paper, we will use the following notational rule. Quantities related to the Bernoulli sieve are starred, whereas corresponding quantities for the general Karlin scheme are not. The same rule applies, for the most part, to perturbed random walks to be defined in Section 3. The numberof occupied boxes intheBernoulli sieve 3 1.2 J - and M -topology: a brief review 1 1 For T >0, let D[0,T] denote the Skorohod space of real-valued functions on [0,T], which are right-continuous with left-hand limits. We will need the J - and the M -topology on D[0,T] 1 1 which are commonly used and were introduced in a famous paper by Skorohod [25]. The J -topology is generated by the metric 1 d(f,g) := inf sup f(λ(t)) g(t) sup λ(t) t , λ∈Λ y∈[0,T]| − |∨t∈[0,T]| − |! where Λ denotes the class of strictly increasing, continuous functions λ : [0,T] [0,T] with → λ(0) = 0 and λ(T) = T. Functions f which are J -convergent to a limit function f are n 1 allowedto have a single jump in the vicinity of a jump of f. Furthermore,the positions of the jumps off andtheir magnitudesshouldconvergetothe positionsofthe jumps off andtheir n magnitude. This is in contrast to locally uniform convergence which requires the positions of jumps of f and f to be the same rather than asymptotically equal. n TheM -topologyisweakerthantheJ -topologyandM -convergenceoff tof isequivalent 1 1 1 n to the convergence of the closed graph of f to the closed graph of f. For instance, choosing n f (t):=1 (t)+2 1 (t) and f(t):=2 1 (t), the f do convergeto f in n [1−1/n,1+1/n) [1+1/n,2] [1,2] n · · theM -topology,butnotintheJ -topologyonD[0,2].Withoutgoingintodetails,wemention 1 1 that the M -topology is typically used in functional limit theorems in which the limit process 1 hasjumpsunmatchedintheconvergentsequenceofprocesses.Thismayhappen,forexample,if the convergingprocessesarea.s.continuousorhaveasymptoticallyvanishingjumps,whilethe limit process is discontinuous with positive probability. We refer the reader to the monograph [26] for a comprehensive exposition of the J - and the M -topologies as well as some other 1 1 topologies on D[0,T]. Throughoutthepaper=J1 and=M1 willmeanweakconvergenceintheSkorohodspacewhen ⇒ ⇒ P endowed with the J -topology and the M -topology,respectively. Furthermore, we will use 1 1 → and also P-lim to denote convergence in probability with respect to P. Finally, =d will stand for equality in distribution. 1.3 Ewens permutations and Ewens sampling formula Let S be the symmetric group of order n. The Ewens family of random permutations is a n parametricfamilyΠ :=Π (θ),θ >0,ofrandomobjectstakingvaluesinS withprobabilities n n n Γ(θ)θ|σ| P Π =σ = , σ S , n n { } Γ(n+θ) ∈ where σ denotesthenumberofcyclesinσ andΓ istheEulergammafunction.Plainly,Π (1) n | | is auniformrandompermutationof 1,...,n forwhichalln!permutationsareequallylikely. { } For r = 1,...,n, denote by C the number of cycles of length r in Π . The following is n,r n the famous Ewens sampling formula: n!Γ(θ) n θci P{Cn,1 =c1,...,Cn,n =cn} = Γ(θ+n) icici!1{Pni=1ici=n}. i=1 Y 4 Gerold Alsmeyer1, Alexander Iksanov2 andAlexander Marynych1,2 Define C (t) := [nt] C for t [0,1]. A remarkable result, originally due to DeLaurentis n r=1 n,r ∈ and Pittel [5] for the uniform case θ =1 and to Hansen [17] for the generalcase θ >0, asserts P that C (t) θtlogn n − =J1 (B(t)) , n , (3) √θlogn ⇒ t∈[0,1] →∞ (cid:18) (cid:19)t∈[0,1] where (B(t)) denotes a standard Brownian motion. Later, much simpler proofs of (3) t∈[0,1] werefoundby Donnelly,KurtzandTavar´e[6]andArratiaandTavar´e[3].While the firstwork is based on a Poisson embedding, the second one uses Feller coupling [2, p. 16] as a key tool. The connection between the Ewens permutations and the Bernoulli sieve emerges when choosing W to have a beta distribution with parameters θ > 0 and 1, i.e. P W dx = { ∈ } θxθ−11 (x)dx. In this case (see, for instance, Example 2 in [14] or Section 5.4 in [2]) (0,1) (K∗ ,...,K∗ ) =d (C ,...,C ), n N n,1 n,n n,1 n,n ∈ which implies that (K∗(t)) =d (C (t)) . (4) n t∈[0,1] n t∈[0,1] 2 Main results and discussion The asymptotic behavior of the small-parts counts in the Bernoulli sieve is well-understood if E( logW ) < . According to Theorem 3.3 in [12], the vector (K∗ ,...,K∗ ), with j fixed, | | ∞ n,1 n,j converges in distribution to a similar vector defined in terms of a limiting “balls-in-boxes” schemeinwhichballweightsareidentifiedwiththearrivaltimesofastandardPoissonprocess on [0, ) and boxes are formed by successive points of exp(A) for a stationary renewal point proces∞s A on R driven by the distribution of logW . A criterion for weak convergence of K∗ | | n is given in Theorem 1.1 and Corollary 1.1 of [10]. One consequence of these results is that, if E( logW ) < , the contribution of the small-parts counts to K∗ becomes asymptotically | | ∞ n negligible, as n , and it raises the natural question which components of the vector (K∗ ,...,K∗ )→prov∞idethemaincontributiontoK∗ forlargen.IfE logW < ,theanswer is pnr,o1vided bny,nProposition 2.1, for the case E logWn= see (13). | | ∞ | | ∞ Proposition 2.1 If µ:=E logW < , then | | ∞ K∗(t) P-lim sup n t = 0. (5) n→∞ t∈[0,1] (cid:12) Kn∗ − (cid:12) (cid:12) (cid:12) (cid:12) (cid:12) Thus, if µ < , the random distribution(cid:12) function t(cid:12) K∗(t)/K∗ converges uniformly in ∞ 7→ n n probabilitytotheuniformdistributionfunction.Thisprovidesadefiniteanswertothequestion above, namely, for each t < s, t,s [0,1], the asymptotic contribution (in probability) of the ∈ vector (K∗ ,...,K∗ ) to K∗ equals s t. n,[nt] n,[ns] n − Letuscomparethisobservationwithsomeresultsofasimilarflavorfromthe literatureand point out beforehand that ρ∗(x), defined by ρ∗(x) := # k N:p∗ 1/x =# k N:W ... W (1 W ) 1/x (6) { ∈ k ≥ } { ∈ 1· · k−1 − k ≥ } for x>0, exhibits a logarithmic growth (see Proposition3.1 below). Consider now the Karlin occupancy scheme with deterministic or random p such that ρ(x), defined by k ρ(x) := # k N:p 1/x , x>0 (7) k { ∈ ≥ } The numberof occupied boxes intheBernoulli sieve 5 is regularly varying at infinity of index α, α (0,1). Let K and K denote the number of n n,r ∈ occupied boxes and the number of boxes containing exactly r balls, respectively, after n balls have been placed. Then lim K /K =c(r) a.s. for explicitly known constants c(r) > 0, n→∞ n,r n seeTheorems8and9in[23],Theorem2.1in[15]andCorollary21in[8](interestingextensions canbe found in[24]).Thus, insharpcontrastto (5), the majorcontributionto K is made by n the small-parts counts. In view of Proposition 2.1, it is natural to ask how fast K∗(t)/K∗ approaches uniformity n n in D[0,1]. We will answer this question by first proving a FLT for the process (K∗(t)) , n t∈[0,1] properlycenteredandnormalized,andthenmakeuseofthecontinuousmappingtheorem.Our mainresults,Theorem2.2andTheorem2.5treatthe caseoffinite andinfinite µ,respectively. Theorem 2.2 Assume that E log(1 W)a < (8) | − | ∞ for some a>0. Set logn u (t) := µ−1 P log(1 W) s ds n {| − |≤ } Z(1−t)logn and logn v (t) := µ−1 P log(1 W) >s ds = µ−1tlogn u (t) n n {| − | } − Z(1−t)logn for t [0,1], where µ=E logW . ∈ | | (A1) If σ2 :=Var logW < , then | | ∞ K∗(t) u (t) n − n =J1 (B(t)) t∈[0,1] µ−3σ2logn! ⇒ t∈[0,1] p as n , and →∞ K∗(t) v (t) tv (1) µσ−2logn n t+ n − n =J1 (B(t) tB(1)) , K∗ − u (1) ⇒ − t∈[0,1] (cid:18) (cid:18) n n (cid:19)(cid:19)t∈[0,1] p where (B(t)) is a standard Brownian motion. t∈[0,1] (A2) If σ2 = and ∞ E(logW)21 ℓ(x), x , {|logW|≤x} ∼ →∞ for some ℓ slowly varying at infinity, then K∗(t) u (t) n − n =J1 (B(t)) µ−3/2c(logn) ⇒ t∈[0,1] (cid:18) (cid:19)t∈[0,1] as n , and →∞ √µlogn Kn∗(t) t+ vn(t)−tvn(1) =J1 (B(t) tB(1)) . c(logn) K∗ − u (1) ⇒ − t∈[0,1] (cid:18) (cid:18) n n (cid:19)(cid:19)t∈[0,1] where c is a positive function satisfying lim c(x)−2xℓ(c(x))=1. x→∞ (A3) If P logW >x x−αℓ(x), x , (9) {| | } ∼ →∞ for some α (1,2) and some ℓ slowly varying at infinity, then ∈ 6 Gerold Alsmeyer1, Alexander Iksanov2 andAlexander Marynych1,2 K∗(t) u (t) n − n =M1 (S (t)) µ−(α+1)/αc(logn) ⇒ α t∈[0,1] (cid:18) (cid:19)t∈[0,1] as n , and →∞ µ1/αlogn K∗(t) v (t) tv (1) n t+ n − n =M1 (S (t) tS (1)) , c(logn) K∗ − u (1) ⇒ α − α t∈[0,1] (cid:18) (cid:18) n n (cid:19)(cid:19)t∈[0,1] where c is a positive function satisfying lim c(x)−αxℓ(c(x))=1 and (S (t)) x→∞ α t∈[0,1] isaspectrallynegativeα-stableL´evyprocesssuchthatS (1)hascharacteristicfunction α u exp uαΓ(1 α)(cos(πα/2)+isin(πα/2)sgn(u)) , u R, (10) 7→ {−| | − } ∈ with Γ being the gamma function. Remark 2.3 Along similar lines as in [10] where weak convergence of K∗ is proved, it can n be checked that moment condition (8) is not needed to ensure weak convergence of the finite- dimensional distributions of (K∗(t)) . Our proof of Theorem 2.2 is based on the decom- n t∈[0,1] position K∗(t) u (t) = K∗(t) ρ∗(n) ρ∗ n(1−t)− n − n n − − (cid:0) + (cid:0)ρ∗(n) ρ∗(cid:0)n(1−t)−(cid:1)(cid:1)(cid:1)un(t) (11) − − withρ∗(x)asdefinedin(6).Itwillbeshownth(cid:0)at,irrespect(cid:0)iveof (8(cid:1)),thefirs(cid:1)ttermontheright- hand side of (11), properly normalized, converges to zero uniformly in probability. However, for dealing with the second term, we need condition (8) (see the proof of Theorem 3.2 below) but do not know whether it is really necessary. Remark 2.4 Whenever the second-orderterm v (t) of the centering of K∗(t) is killed by the n n normalizingconstantswhichisthecase,forinstance,ifE log(1 W) < ,the“true”centering | − | ∞ for K∗(t) is µ−1tlogn. Similarly, the term vn(t)−tvn(1) may then be omitted in the limit the- n un(1) orems for K∗(t)/K∗ because it vanishes asymptotically when multiplied by the corresponding n n normalization. Assume now that W has a beta distribution with parameters θ >0 and 1, giving E logW = θ−1, Var logW = θ−2 and E log(1 W)a < for all a>0. | | | | | − | ∞ Then part (A1) of Theorem 2.2 together with the preceding remark yields K∗(t) θtlogn n − =J1 (B(t)) , n , √θlogn ⇒ t∈[0,1] →∞ (cid:18) (cid:19)t∈[0,1] and in view of (4), this limit relationis equivalent to (3). Thus, we found yet another proof of (3). In fact, our Theorem 2.2 constitutes a generalization of (3). As shown in [11] and [12], respectively, similar generalizations exist for the Erdo¨s-Tu´ran law for the order of Ewens permutations [2, Theorem 5.15 on p. 116] and for the weak laws for small cycles in Ewens permutations [2, Theorem 5.1 on p. 96]. On the other hand, the independence-based tools used in the proofs related to these permutations [2, Section 5] are no longeravailableinthe moregeneralframeworkofthe Bernoullisieveandmustthereforebe replaced by methods of advanced renewal theory. Theorem 2.5 If relation (9) holds with α (0,1), then ∈ The numberof occupied boxes intheBernoulli sieve 7 ℓ(logn)K∗(t) n =J1 W←(1) W←((1 t) ) (12) (logn)α ⇒ α − α − − t∈[0,1] (cid:18) (cid:19)t∈[0,1] (cid:0) (cid:1) as n , and →∞ K∗(t) W←((1 t) ) n =J1 1 α − − , (13) K∗ ⇒ − W←(1) (cid:18) n (cid:19)t∈[0,1] (cid:18) α (cid:19)t∈[0,1] where W←(t):=inf s 0:W (s)>t for t 0 and (W (s)) is an α-stable subordinator (nondecrαeasing L´evy{pr≥ocess) wαith Lapl}ace exp≥onent logαE( zsW≥0 (1))=Γ(1 α)zα, z 0. α − − − ≥ In the remainingpartof this section,we briefly describe our approachandthe organization of the paper. The following heuristic sheds some light on the asymptotic behavior of the number K of n occupiedboxesintheKarlinscheme.Given(pj)j∈N,callaboxk largeifpk 1/n.Onaverage, ≥ a large box k is occupied because the (conditional)mean number ofballs init is np 1.One k ≥ may therefore expect that K is asymptotically close to the number of large boxes, which is n ρ(n) (see (7) for the definition). For the Bernoulli sieve, this heuristic was justified in [10] by showing that K∗ = K∗(1), properly centered and normalized, converges weakly iff ρ∗(n), n n defined in (6) and centered and normalized by the same constants, converges weakly to the samelaw.Fromthis,onemayexpectthat(K∗(t)) iswell-approximatedby(ρ∗(nt)) , n t∈[0,1] t∈[0,1] but this turns out to be wrong. It will actually be shown in Section 4 that the time-reversal ρ∗(n) ρ∗(n(1−t)−) = # k N:n−1 <p∗ nt−1 , t [0,1] − { ∈ k ≤ } ∈ provides a uniform approximation for (K∗(t)) which is tight enough to derive that n t∈[0,1] (K∗(t)) , properly centered and normalized, converges weakly in the Skorohod space if n t∈[0,1] the same is true for (ρ∗(n) ρ∗(n(1−t)−)) . This new observation allows us to replace t∈[0,1] − the existing methods based on Poissonization-de-Poissonization used in earlier works on the Bernoulli sieve and constitutes the first principal contribution of the present paper. We stress that our Lemma 4.1 essentially shows that the behavior of K (t) := [nt] K for large n n k=1 n,r is driven by that of ρ(n) ρ(n(1−t)−) for any Karlin occupancy scheme with deterministic or − P random (pk)k∈N provided that ρ(x) exhibits a logarithmic growth and its increments satisfy an additional condition. Once afunctional limittheoremfor(ρ∗(nt)) has beenproved,the correspondingfunc- t∈[0,1] tionallimittheoremfor(ρ∗(n) ρ∗(n(1−t)−)) followsbyanapplicationofthe continuous t∈[0,1] − mapping theorem. Put T∗ := logW +...+ logW + log(1 W ), n N (14) n | 1| | n−1| | − n | ∈ and note that ρ∗(nt) = # k N : T∗ tlogn equals the number of visits of the perturbed { ∈ k ≤ } random walk (Tn∗)n∈N to the interval [0,tlogn]. Theorem 3.2 stated in Section 3 provides severalfunctionallimit theoremsfor the number ofvisits of ageneralperturbed randomwalk, not necessarily related to the Bernoulli sieve. Being of independent interest, this result is the second principal contribution of the present paper. 3 Perturbed random walks Let (ξk,ηk)k∈N be a sequence of independent copies of a R2-valued random vector (ξ,η) with positive components. Set S :=0, S :=ξ +...+ξ , n N and then 0 n 1 n ∈ T :=S +η , n N. n n−1 n ∈ 8 Gerold Alsmeyer1, Alexander Iksanov2 andAlexander Marynych1,2 The sequence (Tn)n∈N is called a perturbed random walk and has recently attracted some interest in the literature, see [1] and the references therein. Let N(x) denote the number of visits of (Tn)n∈N to the interval [0, x], i.e., N(x) := 1 = 1 , x 0. {Tk≤x} {Sk+ηk+1≤x} ≥ k≥1 k≥0 X X We start with an assertion that will be used in the proof of Proposition 2.1. Proposition 3.1 If m:=Eξ < , then ∞ m(N(n) N(n(1 t) ) lim sup − − − t = 0 a.s. (15) n→∞ t∈[0,1] (cid:12) n − (cid:12) (cid:12) (cid:12) (cid:12) (cid:12) If Eη < , the weak conve(cid:12)rgence of the finite-dimension(cid:12)al distributions of (N(nt)) , t≥0 ∞ properly centered and normalized, follows from Theorem 2.4 in [21], see also Example 3.2 there. Theorem 3.2 given next is an extension of the aforementioned result to convergence in the Skorohod space. The standard approach to such a strengthening would be to prove tightness. Since this turned out beyond our reach we use an alternative approach. Theorem 3.2 Let T >0 and F(x)=P η x ,x 0.In (B1), (B2) and (B3) below, assume further that Eηa < for some a>0. { ≤ } ≥ ∞ (B1) If s2 :=Varξ < , then ∞ N(nt) m−1 ntF(u)du − 0 =J1 (B(t)) (16) √m−3s2n ⇒ t∈[0,T] (cid:18) R (cid:19)t∈[0,T] as n , where m=Eξ < and (B(t)) is a standard Brownian motion. t∈[0,T] →∞ ∞ (B2) If s2 = and ∞ Eξ21 ℓ(x), x {ξ≤x} ∼ →∞ for some ℓ slowly varying at infinity, then N(nt) m−1 ntF(u)du − 0 =J1 (B(t)) (17) m−3/2c(n) ⇒ t∈[0,T] (cid:18) R (cid:19)t∈[0,T] as n , where c is a positive function satisfying lim c(x)−2xℓ(c(x))=1. x→∞ →∞ (B3) If P ξ >x x−αℓ(x), x (18) { } ∼ →∞ for some α (1,2) and some ℓ slowly varying at infinity, then ∈ N(nt) m−1 ntF(u)du − 0 =M1 (S (t)) (19) m−(α+1)/αc(n) ⇒ α t∈[0,T] (cid:18) R (cid:19)t∈[0,T] as n , where (S (t)) is an α-stable L´evy process, S (1) has characteristic α t∈[0,T] α → ∞ function (10), and c is a positive function satisfying lim c(x)−αxℓ(c(x))=1. x→∞ (B4) If (18) holds with α (0,1), then ∈ ℓ(n)N(nt) =J1 (W←(t)) (20) nα ⇒ α t∈[0,T] (cid:18) (cid:19)t∈[0,T] as n , where W←(t) is as defined in Theorem 2.5. →∞ α The numberof occupied boxes intheBernoulli sieve 9 We close this section with two further results that will be needed in our analysis. The first one tells us that the maximal number of visits of a perturbed random walk to subintervals of [0,n+b] of length b grows stochastically more slowly than any positive power of n. Proposition 3.3 For all positive b and c, P-lim n−c sup N(nt+b) N(nt) = 0. (21) n→∞ t∈[0,1] − (cid:0) (cid:1) The second result is a straightforward extension of the fact from renewal theory that the expected number of visits of a random walk with positive increments to intervals of length y is bounded by a linear function in y. Proposition 3.4 For all x,y 0 and appropriate positive C and D, ≥ E N(x+y) N(x) Cy+D. (22) − ≤ (cid:0) (cid:1) 4 The Karlin occupancy scheme: an approximation result In this section, we focus on the Karlin occupancy scheme with deterministic (pk)k∈N. For j =1,...,n, let Z denote the number of balls in the jth box, so that n,j K (t) := 1 , t [0,1], n {Zn,j∈[1,nt]} ∈ j≥1 X gives the number of occupied boxes containing at most [nt] balls. Proposition 4.1 is the first mainingredientto the proofofTheorem2.2.Theconnectionwiththe Bernoullisievebecomes clear when conditioning on (W ). k Proposition 4.1 Let ρ(x) be as defined in (7). Then E sup K (t) ρ(n) ρ n(1−t)− 6 ρ(n) ρ x−1n(logn)−2 n − − ≤ − 0 t∈[0,13]ρ(cid:12)(cid:12)(cid:12)(n) (cid:18)∞ (cid:0) (cid:1)(cid:19)(cid:12)(cid:12)(cid:12) (cid:16) (cid:0) (cid:1)(cid:17) + (cid:12) + t−2(ρ(nt) ρ(n))(cid:12)dt + 2 sup ρ(en1−t) ρ(e−1n1−t) , logn − − Z1 t∈[0,1](cid:16) (cid:17) where x0 >1 denotes an absolute constant that does not depend on n, nor on (pj)j∈N. Proof. Withoutfurther notice,allsubsequentestimates,includingnt n3t/4 >1 fort [l , 1] n with l := 2loglogn, are meant under the proviso that n be sufficien−tly large. We sta∈rt with n logn the basic inequality K (t) ρ(n) ρ n(1−t)− 1 + 1 n − − ≤ {Zn,j>nt,npj∈[1,nt]} {Zn,j=0,npj∈[1,nt]} (cid:12)(cid:12) (cid:0) (cid:0) (cid:1)(cid:1)(cid:12)(cid:12) Xj≥1 Xj≥1 (cid:12) (cid:12) + 1 + 1 {Zn,j∈[1,nt],npj>nt} {Zn,j∈[1,nt],npj<1} j≥1 j≥1 X X =: S(1)(t)+S(2)(t)+S(3)(t)+S(4)(t) n n n n and intend to estimate Esup S(i)(t) for i=1,2,3,4. t∈[0,1] n As for the first summand, we have 10 Gerold Alsmeyer1, Alexander Iksanov2 andAlexander Marynych1,2 sup 1 {Zn,j>nt,npj∈[1,nt]} t∈[0,1] j≥1 X sup 1 + sup 1 ≤ {npj∈[1,nt]} {Zn,j>nt,npj∈[1,nt−n3t/4]} t∈[0,ln)j≥1 t∈[ln,1]j≥1 X X + sup 1 {Zn,j>nt,npj∈(nt−n3t/4,nt]} t∈[ln,1]j≥1 X 1 + sup 1 ≤ {npj∈[1,nln)} {Zn,j−npj>n3t/4,npj∈[1,nt−n3t/4]} j≥1 t∈[ln,1]j≥1 X X + sup 1 {npj∈(nt−n3t/4,nt]} t∈[ln,1]j≥1 X 1 + 1 ≤ {npj∈[1,log2n)} {Zn,j−npj>(npj)3/4,npj∈[1,n]} j≥1 j≥1 X X + sup 1 {npj∈(nt−n3t/4,nt]} t∈[ln,1]j≥1 X =: S(11)+S(12)+S(13). n n n By the definition of ρ, S(11) = ρ(n) ρ(n(logn)−2). n − The random variable Z has a binomial distribution with parameters n and p , so that n,j j EZ =np and VarZ =np (1 p ). Use Chebyshev’s inequality to obtain n,j j n,j j j − np (1 p ) ES(12) j − j 1 (np )−1/21 n ≤ (np )3/2 {npj∈[1,n]} ≤ j {npj∈[1,n]} j j≥1 j≥1 X X = n−1/2x1/2 dρ(x) + n−1/2x1/2 dρ(x) Z[1,n/(logn)2] Z(n/(logn)2,n] ρ(n) + ρ(n) ρ n(logn)−2 . ≤ logn − (cid:16) (cid:0) (cid:1)(cid:17) Finally, the third term S(13) can be estimated as follows: n S(13) = sup 1 n {pj∈(nt−1−n3t/4−1,nt−1]} t∈[ln,1]j≥1 X sup 1 ≤ {pj∈(nt−1(1−n−ln/4),nt−1]} t∈[ln,1]j≥1 X sup 1 ≤ {pj∈(nt−1(1−n−ln/4),nt−1]} t∈[0,1]j≥1 X = sup 1 {pj∈(nt−1(1−(logn)−1/2),nt−1]} t∈[0,1] j≥1 X sup ρ en1−t ρ n1−t . ≤ − t∈[0,1] (cid:16) (cid:0) (cid:1) (cid:0) (cid:1)(cid:17) Summarizing, ρ(n) E sup S(1)(t) 2 ρ(n) ρ n(logn)−2 + + sup ρ en1−t ρ n1−t . n ≤ − logn − t∈[0,1] t∈[0,1] (cid:16) (cid:0) (cid:1)(cid:17) (cid:16) (cid:0) (cid:1) (cid:0) (cid:1)(cid:17) Since S(2)(t) and S(4)(t) are monotone in t, we further infer more easily that n n